# Evaluate an agent with 20-50 tasks from real failures, outcome-based graders and pass^k for consistency

> Start with 20-50 tasks taken from real failures, grade the final outcome in the environment, run several trials per task, and report pass^k when consistency matters.

## Answer

Collect 20-50 tasks drawn from real failures, run each task several times, and score the final state of the environment with graders instead of judging the path the agent took. Report pass@k if one success in k tries is enough, and pass^k if every try must succeed. Checked 2026-09-28 against Anthropic's "Demystifying evals for AI agents".

## Details

### Terms

- **Task:** one test with defined inputs and success criteria.
- **Trial:** one attempt at a task. Run several, because model output varies.
- **Grader:** logic that scores some aspect of a trial. A task can have several.
- **Transcript:** the full record of a trial: outputs, tool calls, reasoning, intermediate results.
- **Outcome:** the final state of the environment, which can differ from what the agent says it did.

### Steps

1. **Collect 20-50 tasks** from real failures. Do not wait for full coverage.
2. **Write an environment reset** so every trial starts from the same state.
3. **Pick graders.** Grade the outcome, not the path, so valid alternative approaches are not penalized.
4. **Run k trials per task** and store every transcript.
5. **Read the transcripts,** especially failures, to check the grader is fair and the task is solvable.
6. **Fix, then rerun.** Keep the suite as a regression check.

### Grader types

| Type | Strengths | Weaknesses |
|---|---|---|
| Code-based | Fast, cheap, objective, reproducible | Brittle to valid variations |
| Model-based | Flexible, captures nuance | Non-deterministic, costs more |
| Human | Gold standard | Slow, costly, needs experts |

### Starter spec template

```yaml
task: "Refund order 1042 if it is within 30 days"
environment_reset: "load fixture orders.json into a fresh test database"
grader:
  type: code
  check: "orders[1042].status == 'refunded' and no other order changed"
k: 5
report: [pass_at_k, pass_hat_k]
```

This layout is our own convention, not a schema from the source.

### pass@k versus pass^k

pass@k is the chance of at least one success in k tries. pass^k is the chance that all k succeed. Assume a per-trial success rate p of 0.8 and independent trials (an assumption, real trials are often correlated). Then pass@k = 1 - (1 - p)^k and pass^k = p^k. Computed with Python:

| k | pass@k | pass^k |
|---|---|---|
| 1 | 0.8 | 0.8 |
| 3 | 0.992 | 0.512 |
| 5 | 0.99968 | 0.32768 |

At k = 5 the agent nearly always succeeds once, yet all five runs succeed only about a third of the time. Use pass^k for customer-facing or unattended work.

### Beyond the eval suite

The source says to combine automated evals with production monitoring, A/B tests and user feedback. For tool design, Anthropic's tools post adds: track runtime, number of tool calls, token use and tool errors, and read raw transcripts to find rough edges.

### Common mistakes

- Grading the exact sequence of tool calls, which fails valid solutions.
- Running one trial per task and treating the result as stable.
- Never reading transcripts, so a broken grader looks like a bad agent.
- Building only synthetic tasks instead of real failures.

## See also

- [[agent-vs-workflow-vs-chatbot]]
- [[when-to-use-multi-agent]]
- [[agent-stuck-in-loop-troubleshooting]]
- [[tool-calling-loop-explained]]

## Sources

- [Anthropic — Demystifying evals for AI agents (published 2026-01-09)](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
- [Anthropic — Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents)

## Sources

- [Anthropic — Demystifying evals for AI agents (published 2026-01-09)](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
- [Anthropic — Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents)