The document you asked the AI to read was also reading the AI back.
Hidden instructions embedded in content that an AI system processes, telling the system to override its own rules: a webpage that, when summarized, instructs the assistant to leak your data; a resume that tells the screening model to rate it highly; an email that, when drafted into a reply, redirects the reply somewhere else. The attack works because the system cannot reliably tell the difference between what you asked it to do and what the content tells it to do. The document is both data and a command. The manipulation is invisible to the person who trusted the tool, because the tool obeyed instructions they never saw.
Truth-adjacency
Truth-independent: the pattern works regardless of whether the claim is true
Where it shows up
Platforms and algorithms
The phrases and tells that mark this pattern in the wild:
an AI assistant behaving strangely after reading a specific document or pagea summary that includes instructions rather than contenta tool that suddenly asks for permissions or actions unrelated to your requestcontent that seems written for a machine rather than a readeran AI output that serves the document's interest rather than yoursThe tell is a split in loyalty you cannot see. A prompt-injected system still sounds like your assistant, but it is partially working for whoever wrote the content it is reading. Watch for outputs that serve the document’s interest rather than your request: a summary that recommends, a draft that redirects, a reply that volunteers information you did not ask to share. The deeper signal is architectural: any tool that reads untrusted content on your behalf can be given orders by that content, and the orders are invisible to you by design. If a tool’s helpfulness suddenly bends toward a stranger’s goal, something in what it read told it to.
You call “prompt injection” on an AI tool that simply gave a bad, biased, or confused answer because of its own training or a hard input. Models err without anyone attacking them. The pattern requires planted instructions: a third party embedded a command in content the system processed, intending to redirect it against your interest. If the odd behavior is the model’s own failure with no external instruction behind it, you are looking at a defect, not an attack. The distinction matters, because one is fixed by better tools and the other is fixed by distrusting what the tool has read.
Field notes where this pattern was identified:
How this pattern gets misused
Someone treats any surprising or unwanted AI output as prompt injection, assuming hidden manipulation whenever the tool does something odd. Most strange outputs are ordinary model errors, not attacks. The term becomes a way to explain away any malfunction as sabotage, which makes the real, demonstrated injections harder to take seriously amid the noise.
What it looks like when you're wrong about it
An AI system producing an odd, unhelpful, or biased output because of its training or a genuine misunderstanding is not under prompt injection. The pattern requires a third party to have embedded instructions in content the system processes, with the intent to redirect the system's behavior against the user's interest. If the strange behavior traces to the model itself rather than to planted instructions in the input, that is a flaw, not an injection.
Not sure? Describe the situation to someone outside it. If they do not see the pattern, pause before you name it.
The name is designed to spread. The hook is designed to stick. If you recognized something, share the name.