Agents loop when completion is ambiguous; cap steps, dedupe tool calls and return explicit terminal states
DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown
Answer
An agent repeats a tool call when nothing in its context tells it the task is finished or that the last call did not help. Put the stop logic in your code, not in the prompt: cap the number of steps, refuse a tool call with identical name and arguments that already ran recently, stop when results stop changing, and return explicit results such as not_found or done. Checked 2026-09-28.
Details
What the vendors document
- Step caps are standard. Anthropic's computer-use reference loop runs
for _ in range(max_iterations)with a default of 10, and the docs say the limit "prevents potential infinite loops that could result in unexpected API costs". - LangGraph reports it as an error.
GRAPH_RECURSION_LIMITis raised when a graph hits its maximum number of steps without a stop condition. The page says that if you do not expect many iterations you probably have a cycle, and that you can raise the limit withgraph.invoke({...}, {"recursion_limit": 1000})only when the graph legitimately needs it. - Errors should be instructive. The Claude docs tell you to say what went wrong and what to try next, instead of "failed". They also say Claude retries an invalid call 2-3 times with corrections.
Symptom, cause and fix
The causes below are this wiki's own advice and practitioner experience, not vendor statements.
| Symptom | Likely cause | Fix |
|---|---|---|
| Same call, same arguments, many times | Tool output gives no sign the call worked | Return an explicit status and the data, for example {"status": "ok", "count": 0} |
| Same tool, slightly different arguments | Empty or vague result, so the model guesses again | Say why it was empty and suggest the next step; cap retries per tool |
| Alternates between two tools | Each tool's output triggers the other | Log the sequence; merge the tools or add a state flag |
| Never says it is done | No completion criterion in the task | State what "done" looks like; add a finish tool or a structured final answer |
| Retries a failing call forever | Error text is generic | Return is_error with a cause and a next action, or tell it to give up |
| Runs until the step limit | Task is too large for one loop | Split the task; see when-to-use-multi-agent |
Guard code
Three rules, all in plain Python. Steps are counted per run. A repeat is the same tool name with the same arguments, compared as sorted JSON so key order does not matter. "No progress" means several identical outputs in a row.
import json
from collections import deque
class LoopGuard:
def __init__(self, max_steps=10, window=6, max_repeats=2, patience=3):
self.max_steps, self.max_repeats, self.patience = max_steps, max_repeats, patience
self.recent = deque(maxlen=window) # (tool, args) keys
self.steps = 0
self.stale = 0
self.last_output = None
def check(self, name, args, output=None):
"""Call after each tool call. Returns None or a reason to stop."""
self.steps += 1
if self.steps > self.max_steps:
return "max_steps"
key = (name, json.dumps(args, sort_keys=True))
self.recent.append(key)
if self.recent.count(key) > self.max_repeats:
return "duplicate_call"
if output is not None and output == self.last_output:
self.stale += 1
if self.stale >= self.patience:
return "no_progress"
else:
self.stale = 0
self.last_output = output
return None
We ran an offline test on 2026-09-28 with Python 3. It checks that a third identical call returns duplicate_call, four identical outputs return no_progress, a step past the limit returns max_steps, and reordered argument keys still count as a repeat. The output was all guard tests passed.
When the guard fires, do not just raise. Return a final tool result such as {"status": "stopped", "reason": "duplicate_call"} with is_error: true, so the model can explain the failure. Every tool call still needs a matching result; see tool-calling-loop-explained.
Common mistakes
- Raising
recursion_limit(ormax_steps) to make the error go away without finding the cycle. - Comparing arguments as raw strings, so
{"a":1,"b":2}and{"b":2,"a":1}look different. - Returning an empty string as a tool result. The model cannot tell "no results" from "broken".
- Putting "do not repeat yourself" in the prompt as the only defence. Prompts are advice, and the guard is enforcement. See instruction-file-ignored-by-agent.
- Not logging each step, so you cannot see which of the three rules should have fired.