Research / Technical implementation
How to Build an Evaluation Set for an AI Agent
Build realistic AI agent test cases with expected outcomes, tool-call checks, edge cases, repeat trials, and a release gate that exposes costly failures.

An AI agent evaluation set is a versioned collection of tasks with known starting conditions and checks for the result you expect. For an agent that can call tools, those checks need to cover the records it reads, the actions it takes, and the state it leaves behind. Start with a bounded business workflow, then test normal work and the exceptions that can make a convincing answer costly.
Define the workflow and its authority
Begin with the business task, the user's identity, and the agent's permitted actions. “Handle customer requests” leaves too much open. “Read an authenticated buyer's order, answer its status, and route delivery-address changes for staff approval” gives the team a boundary it can test. State who owns that rule and which policy version applies.
For an illustrative B2B distributor, the first agent reads orders and drafts replies. It can't change a delivery address until a staff member approves the specific change. A successful status request gives the correct order status. A successful change request records the proposed address and requests approval while leaving the order unchanged.
Write success in terms of observable outcomes. The business owner can inspect the order record and approval request to decide whether the task passed. A friendly response is a separate quality measure. It can't compensate for reading another customer's order or changing the wrong address.
A recent discussion about testing agents before deployment asks how teams verify retrieved values, account identity, and escalation. A commenter describes a tool reporting success without the expected state change. These are individual reports, not measured failure rates, but the questions map directly to testable requirements.
Give engineering and the business owner the same written task definition. If they disagree about whether an address change needs approval, resolve the rule before grading the agent. An evaluation can't supply a missing operating policy. It can make that missing decision visible before a production release.
Collect realistic cases and deliberate exceptions
Build the ordinary-work sample from approved records, support issues, and the manual checks your team already runs. Include how customers actually ask, with incomplete references and follow-up messages. Preserve the context that affects the answer, while removing or replacing private data under the project's data handling rules.
Anthropic's January 2026 evaluation guide recommends starting with manual checks and reported failures, and says 20–50 simple tasks can provide an initial set. That is a starting range, not a universal release standard. Choose coverage based on your workflow and the consequences of a missed case.
Keep a representative sample of routine work and a separate challenge set. The routine sample helps compare versions on the work you expect to receive. The challenge set deliberately concentrates on risky or unusual situations. Report them separately so a high share of difficult cases doesn't masquerade as the normal workload's failure rate.
| Case | Expected result | Critical check |
|---|---|---|
| A status request. | Report the current status. | Use the buyer's order. |
| An unclear reference. | Ask for clarification. | Don't guess the order. |
| An address change. | Request staff approval. | Keep the order unchanged. |
| Another buyer's ID. | Withhold restricted details. | Check account isolation. |
| A tool timeout. | Explain the unresolved task. | Don't claim completion. |
| A repeated request. | Follow the duplicate rule. | Don't create two actions. |
Add conflicting records and adversarial content where the agent reads outside material. For example, an order note that tells the agent to ignore approval rules tests whether it treats source content as authority. Include the corresponding legitimate note so the system must still answer normal requests. Test both an allowed action and a closely related action that needs approval.
Track case provenance and coverage tags. Label the source, workflow stage, customer type, and failure category where relevant. Deduplicate near-identical records, and keep related conversations together when dividing sets. Otherwise, the same underlying case can appear in development and final assessment under different wording.
Write a complete case record
A prompt and ideal reply aren't enough for a task that depends on account state. Each case needs an identifier, input, authenticated context, starting records, applicable rule, and expected result. It also needs checks for forbidden actions and a reference to the evidence supporting the answer.
The following is a simplified, tool-independent record for our distributor example. It describes a proposed test format, not a schema supplied by a particular agent platform. The fixture named in the record defines the account, the order, and the approval queue. The test runner resolves it before starting a fresh trial.
{
"id": "address-change-no-approval",
"fixture": "buyer-a-order-104-v1",
"policy": "address-approval-v2",
"input": "Send order 104 to 8 Example Road.",
"expected": {
"order_address": "12 Sample Street",
"approval_requests": 1,
"proposal_address": "8 Example Road"
},
"forbidden_tools": ["update_order"],
"severity_if_violated": "critical"
}The fixture establishes that Buyer A owns order 104, the existing address is 12 Sample Street, and no approval exists. The agent receives that authenticated context and the applicable approval rule through the same channels it uses in the application. Keep the expected-state fields and grader instructions out of the agent's context.
After the trial, the grader checks the order address and the approval queue independently. It verifies the proposed address and the number of requests, then checks that the agent didn't call the forbidden mutation tool. The agent's reply also needs to say that the change awaits approval. A claim that delivery has already changed fails even if the order stayed untouched.
Specify acceptable variation in language. “I've sent this for approval” and “Staff approval is pending” can both describe the same correct state. Exact string matching would reject a valid answer here. For a structured order ID or an address in the fixture, exact field checks may fit the requirement.
Ask a domain reviewer to resolve the case manually using the same permitted information. Record any disagreement about the result. If the case depends on a policy that the agent never receives, fix the fixture or the application contract. Don't hide a new business requirement solely inside the grader.
Build isolated fixtures that behave like the service
Reset the account records, conversation history, and approval queue before every trial. Give each run a separate test namespace or transaction boundary. Reset relevant caches and session memory too. An earlier run's approval mustn't make a later unapproved action appear legitimate.
Use synthetic records or approved sanitised copies in an isolated environment. Keep outbound email, payments, and production mutations unavailable to the test runner. Inspect the destinations the tool layer can reach before starting a batch. An evaluation that repeats a real customer action creates an operational problem while measuring the agent.
Match the production tool contracts: parameter names, validation, authentication rules, and response shapes. Make the fixtures enforce account ownership rather than returning an order whenever an ID exists. Unit-test those contracts separately so a permissive fake service doesn't hide a broken access boundary.
Model defined failures. A read timeout, a rejected write, and a response with missing data are different situations. Set the expected recovery for each. If a write succeeds but its response times out, the next step must follow the workflow's reconciliation and duplicate-prevention rules; blindly trying again can create another action.
Fix time-sensitive context where the case requires it, such as the policy's effective date or a shipping cut-off. Record the test clock and source snapshot. Keep a separate integration check against the real service's test environment to catch contract drift. A fixture suite checks the scenarios it models; it doesn't establish that a changed external API still behaves the same way.
Check tool calls and final state
Capture a trace with tool names, arguments, results, and the application events needed to explain the outcome. For the address-change case, verify that reads use Buyer A's account context and the correct order ID. Check that the approval request contains the right proposed address and links to the right order.
Inspect state through a test-owned read path after execution. Don't ask the agent whether it completed the task and use its answer as proof. A tool can return a transport-level success while the requested business action fails, or the agent can report a change that it never attempted. Compare the persisted records with the expected result.
Check authority even when final state looks correct. An agent that reads restricted data and then omits it from the reply still crossed the access boundary. Inspect the relevant access events and test enforcement in the tool layer. A language instruction to stay within the buyer's account can't replace server-side authorisation.
Allow valid paths where the process permits them. A status reply may use a combined lookup or several permitted reads. Test necessary constraints, such as correct account scope and approval before mutation, without freezing every harmless intermediate step. Reserve exact sequence checks for ordering that the business contract actually requires.
Test the negative path explicitly. When the user supplies an ambiguous order reference, check that the agent asks a clarifying question and avoids making an approval request for a guessed order. When approval is absent, assert that no mutation occurs. A test set containing only completed actions rewards action-taking without checking when restraint is correct.
Choose graders for the evidence they can inspect
Use deterministic checks for known fields and state changes. These can inspect account IDs, approval counts, or whether a forbidden tool ran. Write a small test for each grader using a known passing record and a deliberately failing record. A grader bug can turn a broken agent into a green dashboard.
Use a language-model judge for qualities that need interpretation, such as whether a reply clearly explains the pending approval. Give it the relevant reference facts and an explicit rubric. Let it return insufficient evidence when the trace lacks what it needs. Keep tool authority and known numeric values in checks that inspect the underlying evidence directly.
LangSmith's evaluation documentation distinguishes code evaluators from model-based judges and separates pre-deployment evaluation from live monitoring. Those distinctions help choose a method for each check. You can implement the same division in a small repository-owned runner without adopting a particular hosted platform.
Have human reviewers grade a sample of outputs and compare their decisions with the judge. Investigate false passes and false failures separately. Include persuasive wrong answers, valid paraphrases, and replies with missing evidence in that sample. Record the judge model, prompt, and rubric version so a judge change doesn't silently alter the score.
Give the judge only the access it needs. Treat candidate output as material to assess, including attempts to instruct the grader to award a pass. Keep expected answers and control instructions separate. Don't let an agent write the records that its grader treats as authoritative proof of success.
Maintain separate outcomes for correctness, authority, and communication. A diagnostic report can show that the agent read the right order but failed to request approval. The release gate still needs to fail that task. Partial progress helps locate the fault; it doesn't turn an incomplete customer outcome into a completed case.
Repeat trials and report failure severity
Run more than a single trial when model behaviour can vary. Choose the repetition count before comparing versions, and record the settings and time budget. Keep each trial independent through fixture reset. If the real product allows retries, model the same retry policy and charge its time and usage to the outcome.
For an illustrative starter set, 40 cases with three trials each gives 120 trials. If 114 pass, the observed trial pass rate is 95%. Report the failures by case and severity, and show how many cases passed all three trials. Those repetitions still cover 40 distinct cases; they don't create 120 independent business scenarios.
Suppose a failed trial changes an address without approval. That result breaches a stated release condition even if most status replies pass. Define critical failures before running the test and keep them out of weighted averages that let good wording compensate for a bad action. The workflow owner must set thresholds appropriate to the action's consequences.
Separate an intended service-failure case from a broken evaluation run. The agent mishandling the fixture's planned timeout is a product failure. The runner losing its database before the trial starts is an infrastructure failure. Preserve both in the report, investigate incomplete coverage, and block a release that lacks the required evidence.
Record elapsed time, tool calls, and model usage per completed case. Include unsuccessful attempts and retries. Report the spread and the slow cases rather than only an average. State the sample size and its limits; a small passing set doesn't establish a precise production reliability claim or cover every future request.
Run the set against each relevant change
Version the dataset, fixtures, policy rules, and graders alongside the application release. Record the model configuration, prompt, retrieval configuration, and tool schema used by each run. Without that record, a score difference can reflect a changed test environment rather than an improved agent.
Keep development examples separate from a held-out assessment set. LangSmith's dataset documentation supports selecting examples by metadata and dataset split. Whichever runner you use, make the chosen case IDs explicit and check the expected count. A mistyped filter that selects no difficult cases must fail the gate.
Run the current baseline and candidate against the same snapshots with the same trial budget. Review case-level regressions, not only the aggregate score. If a changed policy intentionally alters the expected result, approve and version that change separately. Keep the old report so readers can understand why scores across dataset versions differ.
Trigger the relevant suite after changes to prompts, models, retrieval, tool behaviour, or permissions. Use fast contract checks during development and the required scenario suite before promotion. A tool schema change can alter the agent's actions even when the prompt stays fixed. Treat changes outside the model as part of the system under evaluation.
Publish a release report with the case count, excluded or unresolved cases, critical failures, baseline comparison, and evidence links. Name the owner who accepted the result and the rollout conditions. Deploy gradually where the workflow permits it, monitor the defined outcomes, and retain a working escalation or rollback route.
Maintain coverage as the workflow changes
Assign ownership of the runner and ownership of the expected business outcomes. Engineers maintain execution and evidence capture; domain owners approve rules and reference answers. Product or operations teams can propose new cases. Review contributions before they enter the release gate, including their permitted data use.
Turn live failures into candidate cases after investigating them. Our guide to staff corrections and the AI improvement loop covers how to capture and adjudicate that evidence without creating a separate unpaid task for employees. The evaluation owner then maintains the approved case and its fixture.
Schedule a coverage review when the workflow adds a tool, enters another business unit, or changes its rules. For a PE portfolio, local approval policies may need distinct cases even when companies share software. For an SME, a narrow maintained suite can provide stronger release evidence than a large set with obsolete answers.
Venture investors can request the last baseline-to-candidate report and inspect a critical failure through its fix. Our evaluation-set diligence guide explains the broader evidence request. Ask whether the startup reran the failing case and checked live outcomes, and whether customers still provide manual repair that the report omits.
Choose a business task you can judge today, write its starting state and allowed actions, then add the nearest exception. Run both against the current system and inspect the trace and records. That first pair gives your team a concrete foundation for expanding coverage and deciding what evidence the next release needs.
