Too many MCP tools degrade agents; fix it with scoped servers, listChanged and tool search
DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown
Answer
Tool definitions are sent to the model as input tokens, so hundreds of tools crowd out useful context. Anthropic reports that selection accuracy degrades once you exceed 30 to 50 available tools. The fix has three parts: split tools into small domain-scoped servers, keep each server's tool list stable, and load rarely used tools on demand with tool search (defer_loading: true). This page was checked on 2026-09-28 against MCP specification 2026-07-28 and the Anthropic tool search docs. The numbers below are vendor-reported, not measured here.
Details
Why it happens
The Anthropic docs say a multi-server setup (GitHub, Slack, Sentry, Grafana and Splunk) can use about 55k tokens on definitions before any work starts, and that tool search typically cuts this by over 85 percent. MCP itself sets no limit on tool count. The load-everything behavior comes from the client.
Threshold table
| Tools visible to the model | Tactic |
|---|---|
| Under 10, small definitions | Load all. Anthropic says standard tool calling fits better here |
| 10 or more, or definitions above 10k tokens | Consider tool search (Anthropic guidance) |
| About 30-50 and up | Accuracy risk (Anthropic). Split by domain and defer-load |
| 200+ tools across aggregated servers | Tool search or a gateway that filters the visible set |
Steps
- Split by domain. In the community discussion, contributors recommended one server per domain, each with roughly 10 to 30 tools. This is a community pattern, not spec text.
- Write short descriptions. One or two sentences per tool, as another contributor suggested.
- Prefix names when aggregating. The spec says aggregators SHOULD disambiguate colliding tool names, for example by prefixing a server identifier. Anthropic also advises prefixing by service so one search matches a group.
- Keep the list predictable. A server's tool list MAY vary with the caller's authorization but MUST NOT vary per connection. Servers SHOULD keep a deterministic order, which helps caching.
- Signal changes. If the server declares
listChanged, it SHOULD sendnotifications/tools/list_changedto clients that openedsubscriptions/listenwithtoolsListChanged: true. Clients then calltools/listagain. - Defer-load. Add the tool search tool and set
defer_loading: trueon rarely used tools. At least one tool must stay non-deferred, and Anthropic advises keeping your 3 to 5 most-used tools loaded. With the MCP connector, set it once in themcp_toolsetentry'sdefault_config, or per tool inconfigs.
Tool search returns up to 5 tools per search by default, and a request can hold up to 10,000 deferred tools. You still send every definition in each request. Deferral controls what enters the model's context.
Worked token budget
All numbers are illustrative.
- Assume 120 tools at 450 tokens each: 120 × 450 = 54,000 tokens loaded up front.
- Keep 4 tools loaded: 4 × 450 = 1,800. Assume the search tool costs 300.
- One search finds 5 tools: 5 × 450 = 2,250.
- Total: 1,800 + 300 + 2,250 = 4,350 tokens.
- Saving: (54,000 − 4,350) ÷ 54,000 ≈ 92 percent.
A second search adds another 5 × 450 = 2,250 tokens, so the saving shrinks for long tasks.
Common mistakes
- Marking every tool
defer_loading: true. The API returns a 400 error. - Splitting servers but connecting all of them to one agent, which recreates the same list.
- Varying the tool list per connection. The spec forbids it.
- Verbose descriptions and schemas on every tool.
- Trusting a gateway's filtered view without checking what was removed. Filtering is a community pattern and adds its own trust boundary.