Numerical examples are illustrative unless explicitly sourced.
Agent evaluation measures whether the complete system achieves specified goals while respecting constraints and resource limits. The evaluated system includes the model, instructions, tools, permissions, memory and runtime. A fluent final answer is only one piece of evidence.
Remember: Outcome + conduct + consistency + cost.
Learn what success means when software takes actions
For a booking agent, a good final sentence is not the product. The product is the correct reservation, for the correct person, within the allowed constraints, without an unauthorized payment. Evaluation must inspect the resulting world as well as the response.
A task describes the intended outcome and constraints. A trial is one attempt to complete that task. A trajectory is the sequence of observations, decisions, tool calls, and results in the trial. The same task can have several valid trajectories. One agent may search by date and another by destination; neither should fail merely because its path differs from a reference trace.
Define prohibited behavior independently of task completion. A booking that succeeds after exposing another customer's record is not acceptable. Define resource constraints too: completion after an hour and hundreds of calls may be unacceptable for an interactive request. Outcome, conduct, and efficiency are related dimensions, not interchangeable scores.
Build one evaluation case and run it repeatedly
Create an isolated starting state: an account, a set of available reservations, a payment limit, and an explicit instruction. Run the agent, then inspect the reservation and payment records. Confirm the required fields, absence of duplicate bookings, and compliance with the limit. Reset the state before the next trial so the second agent does not inherit the first agent's effects.
Use deterministic checks when the outcome is structured. Use expert or calibrated model grading for aspects such as explanation quality. Inspect traces to diagnose why a trial failed. The model's explanation can help, but it is not proof of its internal reasoning or of successful execution.
Make the evaluation procedure explicit:
- Define the task, starting state, allowed actions and success conditions.
- Pin the system and environment versions.
- Run independent reset trials under a stated budget.
- Grade final state and prohibited behavior separately.
- Inspect failure traces and calibrate ambiguous judgments.
- Report results by task family with costs and uncertainty.
Read diagram source
flowchart LR
T[Versioned task and initial state] --> H[Isolated evaluation harness]
H --> A[Agent trial with bounded tools]
A --> S[Final environment state]
A --> R[Observable trajectory and usage]
S --> G[Outcome and conduct graders]
R --> G
G --> E[Per-task and per-slice results]
E --> D[Diagnose failures and compare releases]
Evaluate changing environments without losing control
Mocked services and snapshots make comparison repeatable, but can omit real latency, interface drift, and failure modes. Live testing captures those conditions but introduces variation and potential side effects. Use both for different purposes: controlled cases for regression and carefully scoped live probes for integration behavior.
For shadow runs, isolate writes and sensitive destinations. For canaries, bound real exposure and verify outcomes. Keep environment versions and reset procedures in the evaluation record. If a test site changed between candidate runs, investigate that difference before attributing the result to model quality.
Long trajectories deserve analysis, but the shortest path is not automatically best. An extra authorization check or clarifying question may be essential. Set step, time, and cost budgets appropriate to task complexity, and examine repeated no-progress behavior. A useful efficiency metric rewards acceptable outcomes within budgets, not reckless speed.
For the booking example, turn “a refundable train ticket under $80” into assertions for reservation existence, traveler, date, refundability, total price, authorization and absence of duplicate or unrelated changes. A grader evaluates some aspect of the trial; a harness runs and records it. This distinction follows Anthropic's agent-evaluation guidance.
Four independent questions
| Dimension | Example measure | What it catches |
|---|---|---|
| Outcome | Verified correct reservations / attempted tasks | Confident claims without real completion |
| Conduct | Unauthorized actions; missed required approvals | Success achieved through unacceptable means |
| Consistency | Repeated success per task and task family | A lucky demonstration |
| Efficiency | Total spend and time per successful outcome | Wasteful loops and expensive retries |
RAG also needs safety, latency, and reliability checks. Agents add state changes and multi-step recovery; the distinction is not “RAG accuracy versus agent safety.”
Pick the grader for the evidence
Use deterministic assertions for balances, schemas, file changes, permissions, and tests. Use expert review for ambiguous policy judgments. Use calibrated model graders for scalable semantic judgments such as whether the response addressed the customer's concern. Inspect false positives and false negatives; a higher judge-model price does not establish accuracy.
Grade observable evidence. Hidden chain-of-thought may be unavailable and is not required for evaluating tool arguments, outputs, approvals, and final state. Protect the grader from instructions embedded in the transcript it is grading.
Do not require a single reference path unless the sequence itself is required by policy. Two valid search strategies can produce the same correct answer. Conversely, an agent that gets the right answer after leaking private data still fails the security gate.
Understand repeated trials
pass@k asks whether at least one of k attempts succeeds. pass^k asks whether all k attempts succeed. Under an illustrative independent identical-success assumption with p = 0.8 and k = 4, these are 1 − 0.2⁴ = 99.84% and 0.8⁴ = 40.96%. Real tasks have varying difficulty and correlated failures; measure repetitions and report the protocol rather than assuming this model fits.
The first metric is relevant when several candidates can be generated and a reliable selector exists. The second reveals consistency. Neither directly proves recovery from perturbations: inject the actual failures you care about, such as timeouts, revoked access, or stale state.
Report denominators, uncertainty and full cost
Suppose 100 attempted tasks produce 80 acceptable outcomes. Model/tool/runtime spending across all 100 attempts, including failed work and retries, is $24. The cost per acceptable outcome is $24 ÷ 80 = $0.30. Dividing only the successful runs' spending by 80 hides failure cost. Add human review and other attributable costs when reporting the product's full unit cost.
| Metric | Define before measuring |
|---|---|
| Task success rate | What counts as an attempted task and an acceptable outcome |
| Tool-call success | Whether transport, schema or business semantics define success |
| Severe violation rate | Severity, denominator and whether one violation fails a release gate |
| End-to-end latency | Start/end events, queue time, human waits and percentiles |
| Cost per acceptable outcome | Failed attempts, retries, tools, infrastructure and review included |
Zero observed failures does not establish a zero failure rate. Under independent, identically distributed Bernoulli trials, zero failures in 300 trials gives a one-sided 95% upper bound of 1 − 0.05^(1/300), approximately 0.994%. This calculation is not valid evidence for untested task families or correlated trials. Report sample size, task mix and observed failures alongside any interval.
Build a useful test suite
Include routine cases, rare consequential cases, impossible requests requiring abstention, ambiguous instructions, malicious tool results, and failures at side-effect boundaries. Version the environment, task data, model, prompt, tools, policy, and graders. Keep development examples separate from release holdouts; inspect and adjudicate broken tests.
Record per-slice results, uncertainty, and paired candidate-versus-baseline differences. A pooled average can hide failure on one language or tenant. Measure tool-call correctness separately from tool availability; many valid calls can still fail the overall task.
Release and ownership
Shadow trials use read-only tools, simulated writes, or isolated cloned state. They must not send real emails or create real orders. Start production exposure within an approved risk envelope; watch completion quality, serious violations, queue load, and costs. Keep a kill switch and a known-good release.
The manager assigns task/rubric ownership to domain experts and product, harness ownership to engineering, adversarial coverage to security, and a named release decision-maker. Evaluation is recurring product work, not a one-time benchmark exercise.
Recall questions
“The new agent uses fewer steps—ship it?” Only if outcome and safety gates still pass; tool calls differ in cost and risk.
“The final answer matches, but a required approval was skipped?” Fail the policy gate even if the outcome is correct.
“The judge and expert disagree?” Review the rubric and evidence, adjudicate examples, measure judge error by slice, and rerun affected evaluations.
Practice by writing five assertions for the ticket example without mentioning a model name. Continue with LLM evaluation.
What familiar agent benchmarks actually test
| Benchmark | Task and environment | What to verify before interpreting a score |
|---|---|---|
| SWE-bench | Resolve repository issues through code changes evaluated in a test environment | Dataset variant, repository snapshots, test harness, tool budget, and contamination; passed tests are not a complete security review |
| WebArena | Complete tasks in reproducible website environments through browser actions | Environment reset, task success evaluator, sites/version, credentials, and whether success reflects the desired final state |
| GAIA | Assistant questions requiring combinations of reasoning, retrieval, tools, or multimodal work | Task level, allowed tools, reference answer, access to external resources, and evaluation protocol |
Do not transfer a repository-fixing score into a claim about safe refunds. Use benchmarks to identify capabilities and failure modes, then evaluate the business workflow's authorized final state, severity, latency, and full cost. Reproducibility includes the agent scaffold and tool versions, not only the base model.
Turn validated traces into the right improvement
Suppose a support agent repeatedly chooses the wrong order when two orders are mentioned. Expert review may show an ambiguous tool schema, missing clarification, or an actual model-selection error. Repair the tool contract or prompt first when that explains the failure. Add the trace as a regression case and compare held-out outcomes.
Training is another possible path, not an automatic consequence of storing traces. SFT requires vetted target behavior. Preference methods such as DPO require meaningful chosen/rejected examples and an appropriate training setup; a raw successful trace and raw failed trace may differ in permissions, difficulty, or tool availability. Split by task/customer/time where needed, remove sensitive data under policy, and check that training improves unseen tasks without weakening action boundaries. Remember trace → diagnose → label → choose intervention → evaluate.
Interview questions with developed answers
Q1: How do you evaluate an agent in a nondeterministic environment such as the web?
Sample answer: I separate controlled regression from live integration testing. Snapshots or mocked services provide repeatable initial conditions and safe writes, while scoped live tests expose interface drift and realistic failures. I verify the required final state and policy compliance, record environment conditions, and repeat trials. I use trajectories for diagnosis rather than requiring an identical reference path when multiple solutions are valid. Any live action is bounded and authorized. A result is interpretable only when I know what changed in both the agent and its environment.
Follow-up: Why reset the environment? Otherwise previous trials can change availability, permissions, or state and bias later results.
Q2: Why is meandering a problem, and how do you address it?
Sample answer: Unnecessary steps can increase cost, delay completion, and expose the agent to more opportunities for error. I inspect traces for repeated requests, unclear tool feedback, missing state, or an objective the agent cannot satisfy. I set task-appropriate limits on steps, time, and cost, plus a policy for no progress. I fix the cause where possible and hand off with preserved evidence when the budget is exhausted. I do not assume every extra step is bad; verification and clarification may be necessary for a correct, authorized outcome.
Follow-up: Is ten steps always enough? No; the budget follows the workload and its acceptable cost and latency.
Q3: What is the difference between task success and tool-call success?
Sample answer: A tool-call success means one operation met its contract, such as returning a valid search result. Task success means the user's intended outcome was achieved under the constraints. Many individually successful calls can still produce the wrong booking. Conversely, a transient tool error can be recovered from and the task can succeed. I measure both, but use authoritative final-state checks for task completion and record policy violations separately. That prevents the agent's own completion message from becoming its grading authority.
Follow-up: Can HTTP 200 count as action success? Only after checking the tool's business result, not just transport status.
Q4: Why report repeated-trial consistency as well as best-of-many success?
Sample answer: A system that solves a task once in several attempts may still be unreliable for a user who receives one attempt. At-least-one success and all-attempts success answer different questions. I report the number of trials, task distribution, reset conditions, and resource budget, and inspect which tasks are unstable. Retry-based product designs must also include the cost and safety of repeated attempts. For side effects, trying again is not free and may require deduplication or reconciliation.
Follow-up: Are repeated trials independent? Often not; shared environment and task difficulty must be considered.
Q5: What must a manager see before increasing agent autonomy?
Sample answer: Evidence of acceptable outcomes, policy compliance, consistency, and operating cost on the intended workload, plus tested containment and recovery. I want to understand the hardest failure cases, review capacity, permissions, and who can stop the system. I would expand one capability or task class at a time with clear gates. A benchmark improvement alone does not justify broader authority, especially when the benchmark omits real tools or human consequences. The decision should connect measured behavior to the actual permission being granted.
Follow-up: What evidence can justify less autonomy? Repeated ambiguous writes, severe policy failures, or review demand beyond capacity.
60-second interview answer
I evaluate the complete agent—model, tools, prompts, permissions, and runtime—on realistic tasks in controlled environments. I check the actual resulting state, not just whether the agent says it succeeded. I also check forbidden actions, approvals, recovery, latency, and total cost. Valid alternative tool sequences should pass. Repeated trials reveal reliability that one successful demo hides. Before rollout, I calibrate model graders against expert labels, isolate shadow side effects, and define release gates for severe failures separately from average task success.