AI on the Clock mehmeterkek.com

Nº 006 · SECURITY · 17 August 2026

What Is the Lethal Trifecta? The Pattern Behind AI Agent Attacks

Private data access, untrusted content, and external communication — each fine alone, dangerous combined. Simon Willison's name for the pattern behind most AI agent security incidents, and how to break it.

The lethal trifecta. Three ordinary capabilities. Combined, they can turn an AI agent into an open door. Manageable alone. Dangerous together.
Three capabilities make up the lethal trifecta, coined by security researcher Simon Willison in June 2025: private data access (it can read your files, email, or database), untrusted content (it processes content controlled by someone else), and external communication (it can send information outside the system). Any one can be useful. The risk emerges when all three meet.
The model can't reliably distinguish data from instructions. An agent may read an email, web page, document, or support ticket containing malicious instructions. Those instructions enter the model's context alongside trusted instructions. The model may then act on them. Prompt injection turns untrusted content into potential instructions.
How the attack runs: Attacker puts malicious content in front of the agent. The content tells the agent to retrieve private information. The agent follows the injected instruction. An external channel carries the data out. No traditional exploit is necessary. The risk comes from the combination of capabilities.
This pattern has been seen in production systems. Researchers have repeatedly demonstrated prompt-injection attacks where untrusted content causes AI systems to access or expose sensitive information. Examples across different products: email with malicious instructions embedded in messages, documents with hidden instructions in files, web content from attacker-controlled pages, and tools or MCP combinations that expose data and provide an exfiltration path.
Break the chain: You don't always need all three. The strongest mitigation is to remove a leg of the trifecta when you can. No private data: limit what the agent can access. No untrusted content: don't expose the agent to attacker-controlled input. No external communication: block outbound channels. If you can remove one leg, the trifecta is broken.
If you can't remove a leg, then constrain all three. Scope data: per-task access, not standing access to everything. Limit tools: expose only the capabilities the task requires. Control output channels: restrict where data can leave. Require approval: put a human or policy gate before consequential actions. Log and trace: make every sensitive tool call reviewable. Reduce what the agent can see, do, and send.
Break one leg. Or constrain all three. Don't rely on the model to recognize the attack. Design the system so the attack can't become an action. Save this before your next agent gets tool access.

Download the deck (PDF, 8 pages)

Private data access, untrusted content, and external communication. Each can be useful on its own. Together, they create what security researcher Simon Willison named the "lethal trifecta" in June 2025.