Skip to content
Wiki

Evaluate an agent with 20-50 tasks from real failures, outcome-based graders and pass^k for consistency

DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown

Answer

Collect 20-50 tasks drawn from real failures, run each task several times, and score the final state of the environment with graders instead of judging the path the agent took. Report pass@k if one success in k tries is enough, and pass^k if every try must succeed. Checked 2026-09-28 against Anthropic's "Demystifying evals for AI agents".

Details

Terms

  • Task: one test with defined inputs and success criteria.
  • Trial: one attempt at a task. Run several, because model output varies.
  • Grader: logic that scores some aspect of a trial. A task can have several.
  • Transcript: the full record of a trial: outputs, tool calls, reasoning, intermediate results.
  • Outcome: the final state of the environment, which can differ from what the agent says it did.

Steps

  1. Collect 20-50 tasks from real failures. Do not wait for full coverage.
  2. Write an environment reset so every trial starts from the same state.
  3. Pick graders. Grade the outcome, not the path, so valid alternative approaches are not penalized.
  4. Run k trials per task and store every transcript.
  5. Read the transcripts, especially failures, to check the grader is fair and the task is solvable.
  6. Fix, then rerun. Keep the suite as a regression check.

Grader types

Type Strengths Weaknesses
Code-based Fast, cheap, objective, reproducible Brittle to valid variations
Model-based Flexible, captures nuance Non-deterministic, costs more
Human Gold standard Slow, costly, needs experts

Starter spec template

task: "Refund order 1042 if it is within 30 days"
environment_reset: "load fixture orders.json into a fresh test database"
grader:
  type: code
  check: "orders[1042].status == 'refunded' and no other order changed"
k: 5
report: [pass_at_k, pass_hat_k]

This layout is our own convention, not a schema from the source.

pass@k versus pass^k

pass@k is the chance of at least one success in k tries. pass^k is the chance that all k succeed. Assume a per-trial success rate p of 0.8 and independent trials (an assumption, real trials are often correlated). Then pass@k = 1 - (1 - p)^k and pass^k = p^k. Computed with Python:

k pass@k pass^k
1 0.8 0.8
3 0.992 0.512
5 0.99968 0.32768

At k = 5 the agent nearly always succeeds once, yet all five runs succeed only about a third of the time. Use pass^k for customer-facing or unattended work.

Beyond the eval suite

The source says to combine automated evals with production monitoring, A/B tests and user feedback. For tool design, Anthropic's tools post adds: track runtime, number of tool calls, token use and tool errors, and read raw transcripts to find rough edges.

Common mistakes

  • Grading the exact sequence of tool calls, which fails valid solutions.
  • Running one trial per task and treating the result as stable.
  • Never reading transcripts, so a broken grader looks like a bad agent.
  • Building only synthetic tasks instead of real failures.

See also

Sources