Skip to content

Concept·Last reviewed 2026-08-18

Prompt injection

Also known as prompt injection attack, indirect prompt injection, hidden prompts.

What is Prompt injection?

Prompt injection is hiding instructions inside content so that an AI system reading that content follows them. Prompt injection works because a model reads text a person may not see, such as white text on a white background or a microscopic font, and cannot reliably tell an instruction from the surrounding material.

How does prompt injection work?

Prompt injection works by placing an instruction where a model will read it and a person will not. A language model processes the text of a document; it does not evaluate whether that text was meant for it, so an instruction embedded in a page can be followed as though it came from the user.

Prompt injection has been documented in the wild. In July 2025 researchers identified 18 manuscripts on arXiv containing instructions hidden in white text and microscopic fonts, aimed at reviewers using AI, from authors at 14 institutions across eight countries.

  • 18 arXiv manuscripts found carrying hidden prompts in July 2025 [1]

Questions this answers

  • How do hidden prompts work?
  • Can text be invisible to humans but readable by AI?
  • Has prompt injection happened in the real world?

Why is prompt injection hard to defend against?

Prompt injection is hard to defend against because a model has no reliable way to separate instructions from data. Both arrive as text in the same context window, and stripping every imperative sentence from an input would remove most legitimate content along with the attack.

Prompt injection is also invisible from the outside. A brand cannot tell from its own analytics that a page an assistant read about it carried an injected instruction, which is why detection has to happen at the answer.

Questions this answers

  • Can prompt injection be prevented?
  • Can you block a prompt injection attack?
  • Why is prompt injection difficult to stop?
  • How would I know if prompt injection affected me?

Key takeaways

  • Prompt injection is hiding instructions inside content so an AI system reading it follows them.
  • Prompt injection works because a model cannot reliably tell an instruction from surrounding data.
  • In July 2025, 18 arXiv manuscripts were found carrying prompts hidden in white text.
  • Prompt injection is invisible from a brand’s own analytics, so it has to be detected at the answer.

Citations

  1. [1] Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review. Lin, arXiv:2507.06185 research, accessed 2026-08-18

Changelog

  1. 2026-08-18 · v1.0 · created. Entity created

Machine-readable chunks for this entity: /wiki/prompt-injection/chunks.json. Corrections to ignite@dhabiai.ae.