The tool-calling loop: send tools, get a tool call, run it yourself, return the result, repeat until the model stops asking
DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown
Answer
Tool calling (also called function calling) means the model never runs anything itself. You send the model a list of tool definitions, it replies with a request to call one of them, your code runs the tool, and you send the output back in the next request. You repeat this until the model replies without a tool request. In the Anthropic API the request is a tool_use block and your reply is a tool_result block that carries the same id. Checked 2026-09-28.
Details
The five steps
- Send the user message plus the tool definitions (name, description, input schema).
- The model replies with a tool call instead of a final answer.
- Your code runs the matching function with the call's input.
- You send a new request that contains the tool output, tied to the call by its id.
- The model answers, or asks for more tool calls. Go back to step 3.
OpenAI's function-calling guide describes the same five steps. The loop lives in your code, not in the model, so you decide when to stop.
Message shapes: Anthropic vs OpenAI
| Anthropic Messages API | OpenAI (Responses API guide) | |
|---|---|---|
| Signal that a call was made | stop_reason is "tool_use" |
function_call items in the output array |
| Call fields | tool_use block: id, name, input |
function_call item with a call_id |
| Result item | tool_result block with tool_use_id, optional content, optional is_error |
function_call_output item referencing the call_id |
| Result role | Inside a user message. There is no tool role |
An item you add to the input list |
| Force or limit calls | See the tool-use docs | tool_choice: auto, required, a specific function, or allowed tools; parallel_tool_calls: false gives zero or one call |
OpenAI's guide also suggests keeping fewer than 20 functions available at once, and calls this a soft suggestion.
Rules the Anthropic API enforces
- The
tool_resultmessage must come directly after the assistant message that holds thetool_use. No message may sit between them. - Inside that user message, all
tool_resultblocks come first. Any text goes after them. Text first causes a 400 error. is_error: truemarks a failed tool. The docs recommend an instructive message, such as what went wrong and what to try next, instead of a bare "failed".- For an invalid call, such as a missing parameter, the docs say Claude retries 2-3 times with corrections before apologizing.
Parallel calls
One assistant turn can hold several tool_use blocks. You may run them concurrently or one after another. You must still return exactly one tool_result per block, all in the next user message. If you skip a call, return a result for it with is_error: true and a short reason.
A runnable loop skeleton
fake_model stands in for the API so the code runs offline. We ran this with Python 3 on 2026-09-28.
MAX_STEPS = 5
TOOLS = {"add": lambda a, b: a + b}
def fake_model(messages):
"""Stands in for the API: asks for one tool call, then answers."""
last = messages[-1]["content"]
if isinstance(last, list) and last and last[0]["type"] == "tool_result":
return {"stop_reason": "end_turn",
"content": [{"type": "text", "text": "2 + 3 = " + last[0]["content"]}]}
return {"stop_reason": "tool_use",
"content": [{"type": "tool_use", "id": "toolu_01", "name": "add",
"input": {"a": 2, "b": 3}}]}
def run(user_text):
messages = [{"role": "user", "content": user_text}]
for step in range(MAX_STEPS):
reply = fake_model(messages)
messages.append({"role": "assistant", "content": reply["content"]})
if reply["stop_reason"] != "tool_use":
return reply["content"][0]["text"], step + 1
results = []
for block in reply["content"]:
if block["type"] != "tool_use":
continue
try:
out, err = str(TOOLS[block["name"]](**block["input"])), False
except Exception as exc:
out, err = f"Error: {exc}", True
results.append({"type": "tool_result", "tool_use_id": block["id"],
"content": out, **({"is_error": True} if err else {})})
messages.append({"role": "user", "content": results})
raise RuntimeError(f"stopped after {MAX_STEPS} steps")
print(run("What is 2 + 3?"))
Output: ('2 + 3 = 5', 2). That is the answer and the number of model calls used. MAX_STEPS makes sure a model that never stops cannot run forever. See agent-stuck-in-loop-troubleshooting.
Common mistakes
- Dropping the result. If the history has a
tool_usewith no matchingtool_resultright after it, the API returns a 400 error: "tool_useids were found withouttool_resultblocks immediately after". Public bug reports show the error in langchain-ai/langgraph issue 5109 and anthropics/claude-code issue 3886. Both are closed. Check that every call id gets a result, including after errors or interrupted runs. - Text before results. Put the
tool_resultblocks first. - Returning only some results when a turn had parallel calls.
- Raising an exception instead of returning
is_error. The model then never learns the tool failed. - No step cap. Always bound the loop.