Multi-agent setups help on broad, parallelizable work and hurt on tightly coupled tasks
DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown
Answer
Use multiple agents when a task is broad, splits into parts that can run independently, and is valuable enough to justify the cost. Use one agent when the steps share context or depend on each other. Anthropic reports that multi-agent systems use about 15x the tokens of a chat, against about 4x for a single agent (vendor-reported, 2025-06-13, checked 2026-09-28).
Details
Terms
A lead agent plans the work and hands parts to sub-agents. Each sub-agent runs its own loop in a clean context window, meaning it starts without the lead's full history. It returns only a condensed summary. Anthropic's context-engineering post describes this pattern and reports summaries of typically 1,000-2,000 tokens.
What the vendor reports
In Anthropic's research system, a lead agent with sub-agents outperformed a single agent by 90.2% on Anthropic's internal research eval (vendor-reported, 2025-06-13; we did not reproduce it). The same post says multi-agent setups fit poorly when agents must share context or have many dependencies. It gives coding as an example, because most coding tasks have fewer truly parallelizable parts than research. It also says the task's value must be high enough to pay for the extra cost.
Decision table
Score the task on three questions.
| Breadth (can parts run independently?) | Coupling (do parts share state?) | Value per task | Choice |
|---|---|---|---|
| Low | Any | Any | One agent |
| High | High | Any | One agent, or a fixed workflow |
| High | Low | Low | One agent; the token multiplier is not worth it |
| High | Low | High | Multi-agent |
Token cost worksheet
The multipliers are vendor-reported and depend on the task. Replace them with your own measurements.
- Measure or estimate the tokens of one chat answer:
chat. - Single agent:
chat x 4. Multi-agent:chat x 15. - Multiply by tasks per month, then by your provider's current token price (look it up, it changes).
Worked example: a chat answer of 10,000 tokens gives about 40,000 tokens for a single agent and about 150,000 for a multi-agent run. That is 3.75 times the single-agent run. If the multi-agent run does not produce a result worth that much more, use one agent.
Checklist before adding agents
- The work splits into parts that need no shared state.
- One agent's context would otherwise fill up with search detail.
- A failed or wrong sub-result can be spotted by the lead or by a grader.
- You have a baseline measurement of one agent on the same tasks. See how-to-evaluate-an-ai-agent.
Common mistakes
- Splitting a coding change across agents that edit the same files, then paying to reconcile the conflicts.
- Assuming more agents means better answers without a single-agent baseline.
- Passing the full history to every sub-agent, which removes the clean-context benefit.
- Ignoring the bill: token use is the main cost driver.