I Hack AI Agent Startups With a Bug From 2011
The Vulnerability in Almost Every AI Agent Nobody's Patching
Pre-Blog Banter
This weekend I flew from SF to Florida for my grandma’s 98th birthday. The whole family came. Cousins I hadn’t seen in years, aunts and uncles from across the country, four generations packed into one living room. Watching her surrounded by all these people she’d shaped over nearly a century, I kept thinking: this is what a life of impact looks like. Dozens of people rearranging their lives just to be in the room with her. I strive to have that kind of impact.
On the flight home, I cracked open my laptop and went back to something I’d been poking at all week. I’d been finding the same vulnerability in AI agent after AI agent. At first glance it looks like the sort of vulnerability that you are use to getting in a generic pentest report from that vendor you have all heard off. Something so boring you wouldn’t even bring it up to the dev team. But in AI Agents … it’s deadly.
I’ll get to what it is. First, let me show you why it’s there.
Every agent has this problem
You’re building an AI agent. It thinks, reasons, calls tools, and responds in a chat interface. The response needs to look good. Markdown, bold text, code blocks, tables. So your frontend renders the agent’s output as HTML, probably using React’s dangerouslySetInnerHTML.

The name is literally a warning. React’s developers named it that on purpose. But when you’re shipping fast, trying to nail the demo before Thursday’s partner meeting, you (or should I say Claude) reach for the thing that works. You’ll fix it later. Later never comes. And now your agent’s chat interface will faithfully render whatever HTML lands in the response body. Whatever HTML. Including the kind you really don’t want.
This is the kind of thing security experts and startup CEOs building fast would scoff and say “So what?” too and usually they would be right. The user would have to inject something malicious into their own conversation. Nobody is going to attack themselves.
They’re right about that part. But they’re missing the actual attack.
The part nobody thinks about
Your AI agent doesn’t just chat. It does things. It reads emails, crawls documents, searches the web, pulls data from integrations. That’s the whole point… It goes out into the world and brings information back.
But what happens when the information it brings back has been poisoned?
An attacker plants a payload somewhere your agent will read. A webpage, a shared document, an email body, a tool response. The agent processes it. Because it’s an obedient little language model doing what the text tells it to do, it includes a malicious snippet in its chat response. Your frontend renders it with dangerouslySetInnerHTML. JavaScript executes in your user’s browser.
Now someone has script execution inside an authenticated session. They can steal tokens, exfiltrate conversation history, grab keys, read anything in the DOM. Your user never typed anything malicious. They just asked their agent to summarize a webpage.
This is called indirect prompt injection, and it turns a bug that everyone dismisses into a full remote exfiltration chain.
I tested this across AI agent products using my own “XSS hunter” Trooper, an easy XSS testing platform. I inject payloads into places the agent reads, wait for the agent to retrieve and render them, watch for callbacks. I kept expecting it to fail. It kept not failing.
The bug
So what’s the vulnerability that’s been solved since 2011, documented in every web security textbook ever written, and is still wide open in most AI agent products shipping today?
It’s XSS. Cross site scripting. The same bug your college professor told you about years ago in college. But you don’t listen.

The AI agent ecosystem just reinvented the conditions for it. Every agent that renders responses as unsanitized HTML is vulnerable. And because agents ingest untrusted content from the open internet by design, the “self” in self XSS stops being a limitation and starts being a delivery mechanism.
Fix it this afternoon
Use DOMPurify. Run every piece of agent output through it before rendering. A few lines of code and the whole chain breaks.
import DOMPurify from 'dompurify';
const cleanHTML = DOMPurify.sanitize(agentResponse);Or stop using dangerouslySetInnerHTML entirely. Libraries like react-markdown render formatted text with builtin sanitization. If your agent’s responses are markdown, and most are, there’s no reason to pipe them through raw HTML.
Or do both. Sanitize server side and client side. Then point Trooperat your agent and see what fires. You might lose an afternoon. You’ll sleep better.
If you’ve raised a Series A, “we’ll fix security later” isn’t a strategy anymore. Your customers are trusting your agent with their data. An unsanitized chat response in a static website is embarrassing, but who cares. An unsanitized chat response in an autonomous agent with access to your customer’s email, documents, and internal tools is how you end up on the news with legal battles with Morgan and Morgan trying to settle.
Ninety eight years of showing up for people, of building something that lasts. That’s what I saw at my grandma’s birthday this weekend. A room full of proof that doing things right compounds over time. Cutting corners doesn’t.
Sanitize your outputs. Test your agents. Fix this before someone else finds it for you.
If you don’t think your agent can be hacked I’m happy to prove you wrong - just reach out to us at hello@pigeonlabs.ai.
Subscribe to Breadcrumbs
New field notes on appsec and AI agent security. Free — unsubscribe anytime.
