Skip to content
Wiki

The tool-calling loop: send tools, get a tool call, run it yourself, return the result, repeat until the model stops asking

DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown

Answer

Tool calling (also called function calling) means the model never runs anything itself. You send the model a list of tool definitions, it replies with a request to call one of them, your code runs the tool, and you send the output back in the next request. You repeat this until the model replies without a tool request. In the Anthropic API the request is a tool_use block and your reply is a tool_result block that carries the same id. Checked 2026-09-28.

Details

The five steps

  1. Send the user message plus the tool definitions (name, description, input schema).
  2. The model replies with a tool call instead of a final answer.
  3. Your code runs the matching function with the call's input.
  4. You send a new request that contains the tool output, tied to the call by its id.
  5. The model answers, or asks for more tool calls. Go back to step 3.

OpenAI's function-calling guide describes the same five steps. The loop lives in your code, not in the model, so you decide when to stop.

Message shapes: Anthropic vs OpenAI

Anthropic Messages API OpenAI (Responses API guide)
Signal that a call was made stop_reason is "tool_use" function_call items in the output array
Call fields tool_use block: id, name, input function_call item with a call_id
Result item tool_result block with tool_use_id, optional content, optional is_error function_call_output item referencing the call_id
Result role Inside a user message. There is no tool role An item you add to the input list
Force or limit calls See the tool-use docs tool_choice: auto, required, a specific function, or allowed tools; parallel_tool_calls: false gives zero or one call

OpenAI's guide also suggests keeping fewer than 20 functions available at once, and calls this a soft suggestion.

Rules the Anthropic API enforces

  • The tool_result message must come directly after the assistant message that holds the tool_use. No message may sit between them.
  • Inside that user message, all tool_result blocks come first. Any text goes after them. Text first causes a 400 error.
  • is_error: true marks a failed tool. The docs recommend an instructive message, such as what went wrong and what to try next, instead of a bare "failed".
  • For an invalid call, such as a missing parameter, the docs say Claude retries 2-3 times with corrections before apologizing.

Parallel calls

One assistant turn can hold several tool_use blocks. You may run them concurrently or one after another. You must still return exactly one tool_result per block, all in the next user message. If you skip a call, return a result for it with is_error: true and a short reason.

A runnable loop skeleton

fake_model stands in for the API so the code runs offline. We ran this with Python 3 on 2026-09-28.

MAX_STEPS = 5
TOOLS = {"add": lambda a, b: a + b}

def fake_model(messages):
    """Stands in for the API: asks for one tool call, then answers."""
    last = messages[-1]["content"]
    if isinstance(last, list) and last and last[0]["type"] == "tool_result":
        return {"stop_reason": "end_turn",
                "content": [{"type": "text", "text": "2 + 3 = " + last[0]["content"]}]}
    return {"stop_reason": "tool_use",
            "content": [{"type": "tool_use", "id": "toolu_01", "name": "add",
                         "input": {"a": 2, "b": 3}}]}

def run(user_text):
    messages = [{"role": "user", "content": user_text}]
    for step in range(MAX_STEPS):
        reply = fake_model(messages)
        messages.append({"role": "assistant", "content": reply["content"]})
        if reply["stop_reason"] != "tool_use":
            return reply["content"][0]["text"], step + 1
        results = []
        for block in reply["content"]:
            if block["type"] != "tool_use":
                continue
            try:
                out, err = str(TOOLS[block["name"]](**block["input"])), False
            except Exception as exc:
                out, err = f"Error: {exc}", True
            results.append({"type": "tool_result", "tool_use_id": block["id"],
                            "content": out, **({"is_error": True} if err else {})})
        messages.append({"role": "user", "content": results})
    raise RuntimeError(f"stopped after {MAX_STEPS} steps")

print(run("What is 2 + 3?"))

Output: ('2 + 3 = 5', 2). That is the answer and the number of model calls used. MAX_STEPS makes sure a model that never stops cannot run forever. See agent-stuck-in-loop-troubleshooting.

Common mistakes

  • Dropping the result. If the history has a tool_use with no matching tool_result right after it, the API returns a 400 error: "tool_use ids were found without tool_result blocks immediately after". Public bug reports show the error in langchain-ai/langgraph issue 5109 and anthropics/claude-code issue 3886. Both are closed. Check that every call id gets a result, including after errors or interrupted runs.
  • Text before results. Put the tool_result blocks first.
  • Returning only some results when a turn had parallel calls.
  • Raising an exception instead of returning is_error. The model then never learns the tool failed.
  • No step cap. Always bound the loop.

See also

Sources