The Quarterly Threat Landscape Report is out. See what attackers are targeting now.Read report

What Is Data Poisoning?

Data poisoning is a cyberattack that corrupts the data an AI or machine learning system learns from or retrieves. It can cause inaccurate outputs, biased behavior, hidden backdoors, or poor security decisions.

Why data poisoning matters

AI systems are only as reliable as the data behind them. So when attackers change that data, they can influence how a model learns, what it retrieves, and how it responds. That makes data poisoning a problem for security, data governance, and business trust.

Data poisoning matters because it can affect systems long before anyone notices an issue. A model may seem to work normally during testing, then fail when it sees a specific phrase, pattern, or type of input. In other cases, poisoned data quietly reduces accuracy over time.

Some of the more common risks include:

  • Unreliable outputs: The model gives incorrect, misleading, or incomplete answers.
  • Biased behavior: Poisoned examples push the model toward skewed or unfair responses.
  • Hidden backdoors: The model behaves normally until a trigger causes a specific failure.
  • Poor security decisions: AI-assisted tools may prioritize the wrong alerts, recommendations, or actions.
  • Loss of trust: Teams may stop relying on AI systems if they cannot explain or validate results.

As more teams use artificial intelligence in security operations, software development, customer workflows, and decision support, the integrity of AI data becomes part of the broader attack surface.

How data poisoning works

Data poisoning usually starts with access to a data source. That source might be a training dataset, a fine-tuning file, a document repository, a knowledge base, or another system the AI uses to learn or retrieve information.

The attack doesn’t always require direct access to the model. In some cases, attackers target the data pipeline around it. A typical data poisoning attack follows this pattern:

  1. The attacker identifies a data source. This could be public web content, an internal file store, labeled training data, or a retrieval-augmented generation (RAG) knowledge base.
  2. They inject or alter data. The attacker adds false examples, changes labels, modifies documents, or inserts misleading content.
  3. The AI system uses the corrupted data. The model trains on it, fine-tunes from it, embeds it, or retrieves it during a response.
  4. The model behavior changes. The system may become less accurate, repeat false information, ignore certain risks, or respond to a hidden trigger.
  5. Security teams investigate the source. Teams may need to trace the output back through prompts, retrieval sources, access logs, and data changes.

This is one reason data poisoning is closely tied to adversarial AI. The attacker isn’t just trying to break a system from the outside, but attempting to shape what the AI system understands as the truth.

Key components of a data poisoning attack

Data poisoning can look different depending on the system, but most attacks involve the same core parts.

Target data source

The target is the data the AI system depends on. For a machine learning model, that might be labeled training data. For a generative AI assistant, it might be a set of internal documents used for retrieval. For a security use case, it could be telemetry, alerts, threat intelligence, or historical incident data.

This is where machine learning and security operations overlap. If teams cannot verify where data came from or how it changed, they may not be able to trust what the model learns from it.

Injection method

The attacker needs a way to get bad data into the system: They may add new records, modify existing files, flip labels, compromise a data feed, or manipulate public content that later gets scraped into a training set.

Some attacks are broad and noisy while others are narrow and precise – designed to affect only one output, one class of inputs, or one trigger condition.

Model learning or retrieval path

The corrupted data needs a path to influence the AI system. That path may be training, fine-tuning, embeddings, or retrieval. In RAG systems, the model may not “learn” the poisoned data permanently, but it can still retrieve and use it in an answer.

This distinction matters for generative AI, as a poisoned document in a connected knowledge base can change what an AI assistant says, even if the underlying model hasn’t been retrained.

Attack objective

The goal may be to degrade accuracy, create bias, hide malicious behavior, or trigger a specific response. In security contexts, attackers may try to make a model ignore certain indicators, misclassify risky activity, or recommend unsafe actions.

Examples and use cases

Training-data poisoning

An attacker adds bad examples to a dataset before a model is trained. For example, they might insert many examples that associate a malicious file pattern with a safe label. If the model learns from those examples, it may classify similar files incorrectly later.

Label flipping

Label flipping changes the “answer key” in a dataset. A model may be trained with examples labeled as safe when they are actually malicious, or labeled as one category when they belong to another.

This can be damaging because the data may look structured and legitimate at first glance. The problem is not the format, rather that the labels teach the model the wrong lesson.

Backdoor trigger

A backdoor poisoning attack trains the model to behave normally most of the time. The failure appears only when the model sees a specific trigger, such as a phrase, token, image pattern, or data condition.

That makes backdoors difficult to detect with ordinary testing. Unless the test includes the trigger, the model may appear accurate.

RAG poisoning

In RAG poisoning, the attacker alters the information an AI system retrieves. For example, they may modify a document in a knowledge base so an AI assistant returns misleading policy guidance, inaccurate security instructions, or false product information.

This is different from prompt injection attacks, where the attacker uses a prompt to manipulate the model at runtime. Data poisoning targets the information the system learns from or retrieves. Prompt injection targets the instructions the model follows during interaction.

How data poisoning fits into security operations

Data poisoning isn’t only an AI engineering issue. It also affects how security teams manage access, monitor changes, validate systems, and respond to incidents.

Security teams should think about AI data pipelines the same way they think about other sensitive systems: who can access them, how changes are approved, what gets logged, and how unusual behavior is investigated.

Important controls usually include:

  • Data provenance: Track where training, fine-tuning, and retrieval data comes from.
  • Access control: Limit who can add, edit, or approve sensitive AI data sources.
  • Change monitoring: Watch for unusual edits, bulk uploads, or unexpected source changes.
  • Validation testing: Test models and AI workflows for accuracy, bias, and trigger-like behavior.
  • Incident response: Preserve affected datasets, logs, prompts, retrieval results, and model outputs for investigation.

AI security posture management (AI-SPM) can help teams frame these controls across AI assets, data sources, users, and connected systems. The goal isn’t to make poisoning impossible, but to reduce the chance that poisoned data enters the pipeline unnoticed and to detect suspicious behavior faster when it’s actually able to enter.

Author

Aaron Wells
Aaron Wells

Frequently asked questions