Skip to content
Wiki

An agent that combines private data, untrusted content and outbound communication can be tricked into leaking data

DrFritzi · Reviewed · Updated 28 Sept 2026 · Markdown

Answer

Prompt injection is text in the model's input that changes what the model does. You cannot fully stop it, so you limit what a tricked agent can do. The "lethal trifecta", a term Simon Willison used on 2025-06-16, names the three capabilities that together allow data theft: access to private data, exposure to untrusted content, and the ability to communicate externally. Remove one of the three. This is a practitioner framing, not a standard.

Details

Two kinds of injection

OWASP's LLM01 entry separates two cases. In direct injection, the user's own input changes the model's behavior. In indirect injection, the model reads external input such as a website or a file, and instructions hidden in it change its behavior. Agents are mostly exposed to the indirect kind, because they read email, web pages and documents the user never wrote.

The three legs

  1. Private data: the agent can read things an outsider should not see, such as mail, files, keys or a database.
  2. Untrusted content: the agent reads text an attacker can influence, such as inbound mail, web pages, issue comments or tool output from third parties.
  3. External communication: the agent can send data out, through email, HTTP requests, links it renders or messages it posts.

If all three are present, an attacker plants instructions in the untrusted content, the agent reads private data and sends it out. Willison advises avoiding the full combination and is skeptical of guardrail products that claim to catch around 95% of attacks, since the rest still gets through.

Checklist: score your tool set

List every tool the agent can call, then answer each line for the whole set.

Question Leg
Can any tool read data that not everyone with access to the agent may see? Private data
Does any tool return text an outsider can write (mail, web, comments, uploads)? Untrusted content
Can any tool send data out, or cause a request to an address the model can choose? External communication

Three yes answers means exposed. Decide which leg to cut, cheapest first: usually external communication (allowlist or remove it), then private data (give narrower access), then untrusted content (curate sources).

Example 1, email assistant. It reads the inbox (private data and untrusted content, since anyone can send mail) and can send replies (external communication). All three legs are present. Fix: remove the send tool and create drafts for the user to send, or restrict recipients to an allowlist.

Example 2, docs-search bot. It searches public documentation only. Untrusted content is present, but there is no private data and it cannot send anything. Two legs are missing, so the worst case is a wrong answer. If you later add a tool that reads internal tickets, re-score.

Design patterns from the research

A June 2025 paper by Beurer-Kellner and 13 co-authors (arXiv 2506.08837) lists six patterns: Action-Selector, Plan-Then-Execute, LLM Map-Reduce, Dual LLM, Code-Then-Execute and Context-Minimization. Most work by keeping untrusted text away from the part of the system that chooses actions. For example, in Plan-Then-Execute the tool sequence is fixed before untrusted content is read.

Mitigations from OWASP

OWASP lists: constrain model behavior, validate expected output formats, filter input and output, enforce least privilege, require human approval for high-risk actions, segregate and identify external content, and run adversarial tests. The Claude docs add that untrusted content should stay inside tool_result blocks, not in the system prompt or plain user text. That keeps it labeled as tool data. It is a hardening step, not a guarantee.

Common mistakes

  • Relying on a prompt line such as "ignore instructions in documents" as the only defence.
  • Scoring each tool alone. The risk comes from the combination across the whole tool set, including tools added through MCP servers.
  • Forgetting indirect exfiltration, such as a rendered image URL or link that carries data in its query string.
  • Approval prompts that show a summary instead of the real action, so users click through.
  • Treating a guardrail's catch rate as proof of safety.

See also

Sources