← Back to Blog

Your AI Agent Evals Need Production Traffic, Not Just Test Cases

Dhaval Bhatt
A glowing AI workflow moving through branching purple and blue production traces while anomalous paths are flagged before release

Your eval suite can pass every test and still miss the failure your next customer finds.

That is not because evals are useless. It is because most early eval suites are built from failures the founder already understands. You write the expected cases, score the outputs, fix the obvious gaps, and ship.

Real users do not follow your test plan.

They bring incomplete inputs, unusual combinations, old documents, ambiguous instructions, and workflows you never expected them to attempt. When you change a model, prompt, tool, or agent harness, the risky question is not only, “Does it still pass our tests?” It is also, “What changes across the work customers actually send us?”

A new release from Raindrop puts that question at the center of AI product testing. Its Simulations product replays production traffic and existing test cases against a proposed agent change, then looks for unexpected behavior shifts before the change reaches users.

The larger lesson matters even if you never use Raindrop. Your production traffic should become part of your evaluation system.

Hand-written evals only test what you thought to ask

A good hand-written eval suite is still the starting point.

It protects the requirements you already know matter. If you are building an AI agent for commercial insurance submissions, you may test whether it:

  • identifies missing documents;
  • extracts the correct policy dates;
  • avoids inventing coverage details;
  • escalates a high-risk exception;
  • produces an output a broker can review quickly.

Those tests encode your domain expertise. Keep them.

But they have a built-in limit. Every case reflects a failure mode someone imagined, observed, or chose to prioritize. The suite can grow over time, but it will always lag behind the messy range of real usage.

OpenAI describes the same coverage problem in its research on deployment simulation. Targeted evaluations remain important for rare, severe, and adversarial risks. Yet they can miss the broader distribution of behavior that appears in normal use. OpenAI’s approach replays realistic conversation contexts with a candidate model to estimate how behavior may change before release.

This is a complement to your fixed eval set, not a replacement.

Your fixed tests defend known requirements. Production-derived tests help expose unknown regressions.

Turn customer work into a living test set

You do not need millions of conversations or a dedicated reliability team to apply this idea.

Start with the workflow your product performs most often. Save a privacy-safe sample of completed runs that represents the actual range of work. Remove personal or confidential details. Preserve the structure that affected the outcome.

Your sample should include:

  • routine jobs that succeed without intervention;
  • jobs users rephrased or retried;
  • runs where a person corrected the output;
  • unusual tool sequences;
  • incomplete or contradictory inputs;
  • cases that triggered a handoff or refund;
  • expensive or unusually slow runs.

Now run the proposed change against that sample in a safe environment. Compare the new version with the current version.

Do not only ask whether the final answer looks good. Check the path the agent took:

  • Did it call more tools?
  • Did it skip a verification step?
  • Did it ask for information the old version inferred correctly?
  • Did cost or completion time increase?
  • Did the escalation rate change?
  • Did a new type of failure appear?

This is where production traffic becomes more than an activity log. It becomes product memory.

Look for behavior changes, not one perfect score

AI agents can fail without producing an obvious error message.

An agent may complete the task but choose the wrong source. It may return a plausible result after a tool failed. It may take eight steps where the previous version took three. It may stop escalating edge cases because a new prompt made it more confident.

These are behavior changes. A single pass rate can hide them.

Raindrop says its simulations combine existing tests with replayed production traffic and anomaly detection. OpenAI reports that deployment simulation helped it find behavior gaps that narrower evaluations missed, while also noting clear limits. Simulation fidelity can introduce error. User behavior changes after a new release. Very rare failures may not appear in the sampled traffic at all.

That is why your release decision should use several signals:

  1. Known-case performance. Did the change preserve the outcomes your fixed evals protect?
  2. Production-sample behavior. What changed across realistic customer work?
  3. Operational measures. Did cost, latency, tool use, or escalation shift?
  4. Human review. Would a domain expert accept the changed behavior?
  5. Staged rollout. Does a small live cohort confirm what the simulation predicted?

No single layer proves the release is safe. Together, they reduce the chance that your customer becomes the test environment.

Your domain expertise decides what counts as a regression

The infrastructure will get easier. The judgment will remain hard.

A general testing tool can show that an agent’s behavior changed. It cannot always tell whether the change matters to your customer.

In healthcare, an extra clarification may be responsible. In a simple scheduling workflow, the same clarification may be friction. In financial operations, a human handoff may protect the customer. In a low-risk content workflow, too many handoffs may destroy the product’s value.

This is the domain-expert founder’s advantage.

You know which shortcuts create risk. You know which exceptions change the decision. You know when a technically correct answer would still fail in practice. That judgment should shape the cases you retain, the anomalies you investigate, and the release bar you set.

Start with twenty hand-written evals. Then let real usage deepen the system. Every correction, retry, escalation, and surprising success can make the next release safer.

The goal is not to build a perfect test laboratory. It is to stop treating production as the first place you learn what your change actually changed.

If you want help turning a workflow you understand into a focused AI product, a reliable evaluation plan, and a 12-week path to launch, book a strategy call. That is the work we do inside AI Product Accelerator.

Sources