For Heads-of · Practitioner

AI Safety

The field of research and practice focused on ensuring AI systems behave as intended, avoid unintended harms, and remain aligned with human values.

  • AI safety
  • alignment
  • safety engineering
  • risk mitigation

Scope

Includes technical safety (preventing unsafe behaviors through training and design), governance safety (policies and oversight), and alignment (ensuring systems pursue intended goals).

Key Areas

Robustness to adversarial attacks, interpretability and transparency, value alignment, specification gaming (systems finding loopholes in their objectives), and scalable oversight.

Governance Context

AI safety is increasingly recognized as essential for responsible AI deployment. Organizations must invest in safety research and practices proportional to the stakes of their AI systems.

← Back to Glossary