For Heads-of · Practitioner

Inadequate fairness testing before deployment

An AI system used in a high-risk domain is deployed without pre-deployment testing for disparate performance or outcomes across demographic groups.

  • medium
  • fairness
  • testing-gap
  • high-risk

How it happens

A system is tested for overall accuracy before launch, but no one runs a subgroup breakdown to check whether that accuracy, or the resulting decisions, hold up evenly across demographic groups.

Why it matters

Aggregate accuracy can hide a system that performs badly, or unevenly, for a specific group, and Annex III of the EU AI Act specifically targets this class of use case for extra scrutiny.

Mitigating controls

The controls that address this risk, ranked by effectiveness.

Framework and clause references

FrameworkClauseTitle
NIST AI Risk Management Framework (AI RMF 1.0)MeasureMeasure
EU AI ActAnnex IIIHigh-risk AI systems

Related resources

The external sources behind this risk, from the Resources library.