For Heads-of · Practitioner
Inadequate fairness testing before deployment
An AI system used in a high-risk domain is deployed without pre-deployment testing for disparate performance or outcomes across demographic groups.
- medium
- fairness
- testing-gap
- high-risk
How it happens
A system is tested for overall accuracy before launch, but no one runs a subgroup breakdown to check whether that accuracy, or the resulting decisions, hold up evenly across demographic groups.
Why it matters
Aggregate accuracy can hide a system that performs badly, or unevenly, for a specific group, and Annex III of the EU AI Act specifically targets this class of use case for extra scrutiny.
Mitigating controls
The controls that address this risk, ranked by effectiveness.
Bias testing and fairness monitoring
Pre-deployment subgroup performance testing plus ongoing production monitoring for disparate outcomes across protected characteristics.
AI impact assessment process
A mandatory, structured assessment of an AI system's impact on individuals, groups, and society, completed before deployment.
Framework and clause references
| Framework | Clause | Title |
|---|---|---|
| NIST AI Risk Management Framework (AI RMF 1.0) | Measure | Measure |
| EU AI Act | Annex III | High-risk AI systems |
Related resources
The external sources behind this risk, from the Resources library.