Skip to content

Concept·Last reviewed 2026-08-18

Data poisoning

Also known as training data poisoning, model poisoning, AI poisoning, backdoor attack.

What is Data poisoning?

Data poisoning is inserting crafted documents into the material a model trains on so as to change its later behaviour. Data poisoning differs from prompt injection in timing: poisoning happens before the model is built, so the altered behaviour ships inside the model rather than arriving with a document it reads.

How much data does poisoning require?

Data poisoning requires far less material than the size of a model suggests. Research published in 2025 by Anthropic, the UK AI Security Institute and the Alan Turing Institute found that roughly 250 malicious documents could plant a backdoor, and that the number stayed near-constant across models from 600 million to 13 billion parameters.

That finding matters because it removes scale as a defence against data poisoning. If the volume of poisoned documents required does not grow with the model, then a larger model is not a safer one on this axis.

  • ~250 documents needed to backdoor models from 600M to 13B parameters [1]

Questions this answers

  • How many documents does it take to poison a model?
  • Are bigger AI models harder to poison?
  • Is data poisoning practical?

Can safety training remove a poisoned backdoor?

Safety training does not reliably undo data poisoning. The same body of research reports that a backdoor can persist after a poisoned model undergoes later safety training, which means the standard mitigation is not a guaranteed one.

Data poisoning is therefore a supply-chain problem rather than a filtering problem. Once poisoned material is in the training set, catching it afterwards is materially harder than keeping it out.

Questions this answers

  • Can you fix a poisoned model?
  • Does safety training remove backdoors?
  • Is data poisoning permanent?

Key takeaways

  • Data poisoning is inserting crafted documents into training data to change a model’s later behaviour.
  • Research published in 2025 found data poisoning with roughly 250 documents could backdoor models from 600M to 13B parameters.
  • The documents required for data poisoning stay near-constant in number as models grow, so scale is not a defence.
  • A data poisoning backdoor can survive later safety training.

Citations

  1. [1] Poisoning attacks on LLMs require a near-constant number of poison samples. Anthropic, UK AI Security Institute, Alan Turing Institute research, accessed 2026-08-18

Changelog

  1. 2026-08-18 · v1.0 · created. Entity created

Machine-readable chunks for this entity: /wiki/data-poisoning/chunks.json. Corrections to ignite@dhabiai.ae.