For Heads-of · Practitioner
Training data provenance checks
Documented verification of the source, licensing, and integrity of any dataset before it is used for training or fine-tuning.
- preventive
- data-provenance
- poisoning
- supply-chain
What it does
Requires a documented provenance check, source, licence, collection method, known contamination, before a dataset is approved for training or fine-tuning use.
Where it fits
The upstream control for both data poisoning and IP/licensing risk; catches a bad dataset before it ever reaches a training run.
Risks this mitigates
The risks this control addresses, ranked by effectiveness.
Data and model poisoning
Training, fine-tuning, or embedding data is deliberately manipulated to introduce vulnerabilities, backdoors, or bias into a model.
LLM application supply chain vulnerabilities
Vulnerable or unvetted components — base models, adapters, datasets, plugins, deployment platforms — enter an LLM application through its supply chain.
Framework and clause references
| Framework | Clause | Title |
|---|---|---|
| NIST Generative AI Profile (NIST AI 600-1) | Intellectual Property | Intellectual Property |
| MITRE ATLAS | AML.T0020 | Poison Training Data |