Skip to content
Back to Insights
Agentic AIBy KE Engineering Team

Evals for Production Agents

Evals for Production AgentsAgentic AI cover for Evals for Production AgentsL4 SHADOWL3 REPLAYL2 UNITL1 REGRESSIONAGENTIC AIEvals for ProductionAgents// FOUR LAYERS · ONE LOOP

Most AI eval content is about static prompt regression testing. That works for stateless LLM calls. It stops working once your system takes multiple steps, uses tools, and has to decide when to stop.

Agent evals are different. Here's the loop we run for every agentic system we ship, and the specific things static eval frameworks miss.

The principle: score the path, and the answer

The point of agent evals is to score the whole path the agent took, and whether the final answer was right. A correct answer reached through a wrong or unsafe path is still a failure. It's a failure you got away with once. Hold that principle and the rest of the eval design follows.

Three failure modes static evals rarely catch

Trajectory failures. The final answer is right, but the path to it is wrong. Agent called seven tools when two would have worked. Or called the tools in an order that leaked data to a downstream system. Output-only evals give little signal on this.

Termination failures. Agent knew the answer at step 3 but kept going until step 12. Or stopped at step 4 when the answer was in step 5. Termination is a policy the model has to learn, and it drifts with prompt changes. Static evals miss it by construction, because they only see the last message.

Tool-selection failures. Agent picked the wrong tool from a list of similar ones. Selected get_customer_orders when it should have used get_customer_orders_recent. The output looks plausible, the eval passes, the wrong data flows downstream. The eval catches it; the fix lives in the catalog, as Tool Design Is Prompt Design covers.

Output-only suites weren't built to catch any of these. You need evals that see the steps.

The four-layer eval architecture

Four-layer eval architecture stackFour horizontal layers stacked from Layer 1 at the bottom to Layer 4 at the top. Left annotation shows production signal increasing upward; right annotation shows cost per run increasing upward.EVAL LAYERSproduction signal ↑cost per run ↑LAYER 04Production shadow evalsproduction traffic. compare live to canary.LAYER 03Trajectory evalsmulti-step scenarios. score the whole path.LAYER 02Tool-choice unit evalsgiven state X, which tool wins?LAYER 01Prompt regression evalsstatic input/output pairs. fast, cheap.// SIGNAL AND COST RISE TOGETHER
Fig. 1: Static tests catch strings. Production quality lives higher up.

Layer 1 is static prompt regression: fixed inputs, expected strings, run on every change. Most teams have it, and it is the floor. Layer 2 scores single tool-choice decisions in isolation. Layer 3 scores whole trajectories. Layer 4 runs a canary against production traffic. Layers 2, 3, and 4 are where production agent quality lives.

Layer 2: tool-choice unit evals

Isolate one decision: given this state and this tool catalog, which tool should the agent pick? Score with an LLM-as-judge or a deterministic rule.

ts
const toolChoiceCase = {  state: {    conversation: "user asked for their most recent 5 orders",    available_tools: ["get_customer_orders", "get_customer_orders_recent", "search_orders"]  },  expected: "get_customer_orders_recent",  rationale: "recent-specific tool exists and matches the query shape"};

Run these on every prompt change. They catch the similar-tool failure mode that end-to-end evals miss.

Layer 3: trajectory evals

Run production multi-step scenarios. Score the whole trajectory: step count, tool ordering, budget consumed, final output correctness, and an LLM-as-judge score for whether a senior engineer would approve the path. This is where the score-the-path principle becomes machinery. A run that lands the right answer after touching a system it had no business touching should fail here, loudly.

Store trajectories as structured logs you can query:

ts
type TrajectoryStep = {  step: number;  tool: string;  arguments: Record<string, unknown>;  result_summary: string;  latency_ms: number;  cost_usd: number;};

Trajectory evals are expensive to run. Run them nightly on a stable set, and on every prompt change to a small critical set. They are also the most useful diagnostic when something breaks in production. The shape of the path being scored: Durable Agents.

Layer 4: production shadow evals

The best signal you'll ever get: production traffic. Route a small percentage of production requests through a canary agent (new prompt, new model, new tool schema) alongside the live one. Compare trajectories. Alert on divergence rate above a threshold.

Two rules for shadow evals to be safe:

The canary is read-only for side effects. Any tool that writes to a production system is mocked or short-circuited. The canary produces trajectories only.

Divergence is measured, not judged. Different trajectories aren't automatically bad. Score them, compare cost and correctness, decide from the aggregate. A canary that always agrees with production isn't teaching you anything.

The loop itself

Evals aren't a test suite you write once and check off. They are a cycle you run continuously, against a baseline that moves as the system improves.

  1. Run. Execute the eval set on a schedule and on every prompt, model, or tool-schema change.
  2. Compare to baseline. Every score is relative to the last accepted run.
  3. Flag divergence. A regression is a move away from baseline larger than run-to-run noise, per layer and per action class.
  4. Triage. Decide whether the divergence is a production regression, an eval defect, or a legitimate behavior change.
  5. Fix. Change the prompt, the tool catalog, the routing, or the eval case that was wrong.
  6. Re-baseline. Accept the new run as the reference and start again.

The loop feeds itself. Production failures become new eval cases. New eval cases tighten the baseline. A tighter baseline surfaces smaller regressions earlier. That compounding is the whole value; a one-time eval run gives you a number, a loop gives you a trend.

Why this is a data problem

Running evals as a continuous loop, holding a baseline, and detecting divergence over time is an observability and data problem as much as an AI problem. Trajectories are event streams. Baselines are versioned datasets. Divergence detection is time-series work with the usual questions about windows, seasonality, and noise. It is the same event-streaming, metrics, and pipeline discipline as any real-time data system, pointed at a newer workload. Teams with that foundation build better agent evals, because the prompt was never the hardest part.

Approval rate and standing review

Before an agent has autonomy, a human reviews every proposed action. Track the rate at which a human approves the action unchanged, per action class, over time; that number is the evidence for graduating a class from proposed to autonomous. For anything irreversible or safety-relevant, review is a standing control, and for some classes it never ends. The Approval Gate covers how to design that review so it keeps working.

What to alarm on

Static prompt regression pass rate below threshold. Tool-choice accuracy dropping. Trajectory length distribution shifting. Cost per successful trajectory rising. Approval rate falling on any action class currently autonomous. First-decision latency degrading (agent is thrashing before picking its first tool).

If cost is the pressure that sends you looking at a cheaper model, this loop is what makes that switch safe; LLM Cost Control Engineering covers the cost side.

Alarm on the loop, and the endpoint. By the time output quality degrades measurably, the trajectory has been rotting for a while.