For Heads-of · Practitioner

Model backdoors and trojans

A hidden, trigger-activated backdoor is embedded in a model via poisoned training data, weight tampering, or a payload injected into the model artifact.

  • critical
  • adversarial-ml
  • backdoor
  • mitre-atlas

How it happens

A model is trained, fine-tuned, or shipped with a hidden trigger, an unusual phrase, an image watermark, a specific input pattern, that causes it to produce an adversary-chosen output on demand, while behaving normally otherwise.

Why it matters

A backdoored model passes every ordinary evaluation, since the malicious behaviour only appears for the trigger the adversary knows and no one else is testing for.

Mitigating controls

The controls that address this risk, ranked by effectiveness.

Framework and clause references

FrameworkClauseTitle
MITRE ATLASAML.T0018Backdoor ML Model

Related resources

The external sources behind this risk, from the Resources library.