Evaluate an agent with 20-50 tasks from real failures, outcome-based graders and pass^k for consistency
DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown
Answer
Collect 20-50 tasks drawn from real failures, run each task several times, and score the final state of the environment with graders instead of judging the path the agent took. Report pass@k if one success in k tries is enough, and pass^k if every try must succeed. Checked 2026-09-28 against Anthropic's "Demystifying evals for AI agents".
Details
Terms
- Task: one test with defined inputs and success criteria.
- Trial: one attempt at a task. Run several, because model output varies.
- Grader: logic that scores some aspect of a trial. A task can have several.
- Transcript: the full record of a trial: outputs, tool calls, reasoning, intermediate results.
- Outcome: the final state of the environment, which can differ from what the agent says it did.
Steps
- Collect 20-50 tasks from real failures. Do not wait for full coverage.
- Write an environment reset so every trial starts from the same state.
- Pick graders. Grade the outcome, not the path, so valid alternative approaches are not penalized.
- Run k trials per task and store every transcript.
- Read the transcripts, especially failures, to check the grader is fair and the task is solvable.
- Fix, then rerun. Keep the suite as a regression check.
Grader types
| Type | Strengths | Weaknesses |
|---|---|---|
| Code-based | Fast, cheap, objective, reproducible | Brittle to valid variations |
| Model-based | Flexible, captures nuance | Non-deterministic, costs more |
| Human | Gold standard | Slow, costly, needs experts |
Starter spec template
task: "Refund order 1042 if it is within 30 days"
environment_reset: "load fixture orders.json into a fresh test database"
grader:
type: code
check: "orders[1042].status == 'refunded' and no other order changed"
k: 5
report: [pass_at_k, pass_hat_k]
This layout is our own convention, not a schema from the source.
pass@k versus pass^k
pass@k is the chance of at least one success in k tries. pass^k is the chance that all k succeed. Assume a per-trial success rate p of 0.8 and independent trials (an assumption, real trials are often correlated). Then pass@k = 1 - (1 - p)^k and pass^k = p^k. Computed with Python:
| k | pass@k | pass^k |
|---|---|---|
| 1 | 0.8 | 0.8 |
| 3 | 0.992 | 0.512 |
| 5 | 0.99968 | 0.32768 |
At k = 5 the agent nearly always succeeds once, yet all five runs succeed only about a third of the time. Use pass^k for customer-facing or unattended work.
Beyond the eval suite
The source says to combine automated evals with production monitoring, A/B tests and user feedback. For tool design, Anthropic's tools post adds: track runtime, number of tool calls, token use and tool errors, and read raw transcripts to find rough edges.
Common mistakes
- Grading the exact sequence of tool calls, which fails valid solutions.
- Running one trial per task and treating the result as stable.
- Never reading transcripts, so a broken grader looks like a bad agent.
- Building only synthetic tasks instead of real failures.