In one paragraph

A risk appetite statement earns its place by settling arguments before they happen. When a model's error rate climbs, nobody should be debating whether that is acceptable, because the answer was agreed in advance, in writing, by the board. Most AI risk appetite statements fail that test, and they fail predictably. They express appetite as a single qualitative posture across all AI activity. Nothing can falsify it, so nothing ever breaches it, so it never causes anything to happen. The fix is structural and not editorial. Appetite for AI has to be set across several distinct risk dimensions, at several tiers of decision consequence, expressed in metrics the organisation actually instruments, with a pre-agreed action attached to each threshold. That is more work than a paragraph in the annual report. It is also what separates a statement the board approves from a statement the board can govern with.

Why the enterprise statement does not survive contact with AI

Enterprise risk appetite statements work because the risks they describe behave consistently. Credit risk in a lending book behaves the same way in January and July, and the appetite that governed it last year governs it this year.

AI risk has three properties that break this.

The first is that AI risk is not one risk. A system can fail by being wrong, by being unfair, by being inexplicable, by being manipulated, by acting without authority, or by depending on a supplier that changes underneath you. Those are different exposures with different owners, different metrics, and genuinely different appetites. An organisation might reasonably accept a meaningful error rate in a document-tagging system while accepting none at all in that same system's handling of protected characteristics. Collapse them into one posture and the information is gone.

The second is that appetite depends on consequence and not on the system. The same model, at the same accuracy, is prudent in one deployment and reckless in another. A 4% false-positive rate is unremarkable where a human reviews every flagged item, and unacceptable where applications are declined automatically. So appetite cannot attach to the model. It attaches to the decision the model influences.

The third is that it degrades while nobody acts. A credit policy does not become more permissive because time passed. A model does. Model drift means the risk position on the day of the board's approval is not the risk position six months later, and no decision was taken by anyone. An appetite statement with no re-measurement cadence is a snapshot presented as a control.

Together these mean the useful unit is a matrix, not a statement: risk dimension against decision-consequence tier, with a measurable threshold in each cell.

Appetite, tolerance, capacity, and the one that gets left out

Three terms are routinely used interchangeably. The confusion is not pedantic, because it determines who is allowed to decide what.

Appetite is how much risk the organisation chooses to accept in pursuit of its objectives. Forward-looking, strategic, set by the board. A statement of intent.

Tolerance is the boundary of acceptable variation around that intent, expressed as measurable thresholds. It is operational, and management sets it within the appetite the board has approved. Appetite says "we accept limited error in customer-facing automation." Tolerance says "false-positive rate above 3.5%, measured weekly."

Capacity is the maximum exposure the organisation could absorb before something breaks: regulatory standing, operational resilience, solvency. It is a fact about the organisation and not a choice. Appetite should sit meaningfully below it. Where the two sit close together, the organisation is running without margin whatever its statement says. The risk appetite glossary entry draws this distinction, and it is the one most often lost in drafting.

There is a fourth element, and its absence explains more failed appetite statements than anything else.

A trigger is what happens when a tolerance is breached, decided in advance. "Escalate to the risk committee" does not qualify, because that describes a meeting. A trigger is the specific thing that occurs: the system reverts to human review, the feature is disabled for new users, the model rolls back to the previous version, deployment to the next segment pauses. Without one, a threshold breach produces a discussion. Discussions held during an incident are how organisations find out their appetite statement was decorative.

The six dimensions an AI appetite has to cover

This is the working set. Use fewer than six and distinct exposures get collapsed together. Use many more and the matrix stops being something a board can hold in mind.

Performance and accuracy. Appetite for the model being wrong, expressed as error-rate ceilings. The shape of the error matters as much as the rate, because false positives and false negatives usually carry very different consequences, and a single accuracy figure hides that.

Fairness and discrimination. Appetite for disparate outcomes across protected characteristics. Boards frequently want to write "zero", which is unmeasurable and therefore untestable. The usable form is a maximum disparity ratio between groups, on a named metric, with a named measurement population, and a genuine commitment that breaching it stops deployment.

Transparency and explainability. Appetite for deploying systems whose outputs cannot be explained to the person affected, or to a regulator. Procurement often sets this one implicitly. A decision to buy a closed model is a decision about explainability appetite, and the board rarely sees it framed that way.

Security and adversarial exposure. Appetite for the system being manipulated: prompt injection, training data poisoning, model extraction, exfiltration of information through outputs. Distinguish appetite for exposure from appetite for unmitigated exposure. The second is what the threshold should measure.

Autonomy and human oversight. How much decision authority is delegated without a human able to intervene. Organisations most often fail to write this one down, and it is the dimension that most reliably produces the incident, because autonomy expands incrementally. A system that recommends becomes a system that pre-fills becomes a system that decides, and there is no single decision to approve along the way. Setting appetite here means naming the human oversight model required at each consequence tier, in advance of the pressure to remove it.

Third-party and provenance. Appetite for depending on models, data or infrastructure the organisation did not build and cannot inspect. Include concentration. Two suppliers built on the same underlying foundation model are one dependency wearing two coats. See third-party risk for the broader exposure this sits inside.

Consequence tiers, and why the matrix has two axes

Against those six dimensions, set three or four tiers describing what the AI system's output actually does. A workable set runs as follows.

Tier 1 is informational. Output is presented to a person who decides independently: search, summarisation, drafting, tagging.

Tier 2 is advisory. Output materially shapes a human decision and in practice is usually followed, covering recommendations, prioritisation, and risk scoring under review.

Tier 3 is automated. Output takes effect without a human in the loop for each instance: approvals, routing, pricing, content publication.

Tier 4 is consequential and automated. The decision materially affects a person's rights, finances, health, employment or access to services.

Appetite falls sharply across the tiers, and it should. Tier 4 is where most regulatory obligation concentrates, and where appetite should approach zero for several dimensions.

The value of the two-axis form is that it turns arguments into lookups. A team proposing to move a system from Tier 2 to Tier 3 is not making a product decision to be negotiated. They are proposing to operate against a different set of thresholds, which the board has already set. That is what an appetite statement is for.

A worked statement

Illustrative, for a mid-sized regulated services firm, abbreviated to three dimensions. The full six-dimension matrix is in the accompanying template.

AI Risk Appetite Statement (extract)
Performance and accuracy. We accept measurable error in AI-assisted processes where a human retains the decision. We have no appetite for automated decisions whose error rate is not continuously measured against a defined baseline. Tolerance: Tier 1, no fixed ceiling, quality reviewed quarterly. Tier 2, precision ≥ 0.85 on the validation set, measured monthly. Tier 3, precision ≥ 0.95 and false-negative rate ≤ 2%, measured weekly. Tier 4, not permitted without Board approval of a system-specific threshold. Trigger: two consecutive measurement periods below threshold, or any single period more than 20% below, reverts the system to the next lower tier until remediated and re-approved by the AI governance forum.
Fairness. We have no appetite for AI systems producing materially different outcomes across protected characteristics in Tier 3 or Tier 4 deployments. Tolerance: maximum outcome disparity ratio of 1.10 between any two groups of ≥100 in the measurement population, on the approved fairness metric, measured at each release and quarterly in production. Trigger: breach halts new deployment immediately and initiates review within five working days. Existing deployment continues only on written risk acceptance by the accountable executive, time-limited to 90 days.
Autonomy and human oversight. We accept automation of decisions at Tier 3 where oversight is designed in and evidenced. We have no appetite for Tier 4 automation without a documented, tested human intervention path. Tolerance: every Tier 3 and Tier 4 system maintains a human override exercisable within one business day, tested twice yearly. Systems escalating one tier require re-approval before the change takes effect. Trigger: a failed override test moves the system to Tier 2 within ten working days.

Those figures are illustrative. Do not adopt them without deciding they are right for your organisation.

Three properties distinguish this from the statements that do not work. Every threshold names a metric, a population and a measurement frequency. Every dimension has a trigger that is an action and not a meeting. And the Tier 4 row says "not permitted without specific approval", which is a real answer. Appetite statements are allowed to say no.

What the frameworks actually require

Useful to know, because appetite is frequently treated as optional good practice when for some organisations it is an explicit obligation.

ISO/IEC 42001, Clause 6.1.2 requires the organisation to establish and maintain AI risk criteria, including criteria for accepting risk. That is the standard's formal expression of appetite. Clause 6.1.3 then requires each identified risk to be treated by a documented choice to mitigate, avoid, transfer or accept, with the reasoning captured in a Statement of Applicability. Clause 6.1.4 adds the AI-specific requirement for an impact assessment covering consequences for individuals, groups and society. In practice, an auditor asking how acceptance decisions were made is asking to see appetite. The ISO 42001 crosswalk shows how this maps against the other frameworks.

NIST AI RMF, GOVERN 1.3 requires that "processes and procedures are in place to determine the needed level of risk management activities based on the organization's risk tolerance." Its playbook is explicit that tolerance should account for distinct sources of risk, covering financial, operational, safety and wellbeing, reputational and model risk. The framework deliberately declines to prescribe levels. It requires that the organisation set them, and apply them consistently across its AI portfolio.

PRA SS1/23 is not advisory for UK banks. Principle 2 states that "the board should set a model risk appetite that articulates the level and types of model risk the firm is willing to accept", and requires model performance to be monitored against a board-approved appetite, with aggregate model risk kept within it. The supervisory statement took effect on 17 May 2024. For in-scope firms, a model risk appetite is a supervisory expectation with a date attached, and every material AI system is a model. That makes it the sharpest version of the obligation in UK regulation, and worth reading directly if you are in scope. See financial services.

Three failure modes worth naming

The unfalsifiable appetite. "We have a low appetite for AI risk." Nothing can breach it, so nothing ever does, so it never changes a decision. Test: can you name a specific measurement that would put the organisation outside its stated appetite? If not, you have written a value.

The uninstrumented threshold. Appetite set against metrics nobody collects. Fairness thresholds defined on a protected-characteristic breakdown the organisation does not record. Drift thresholds with no production monitoring behind them. Test: for each threshold, name the system that produces the number and the person who reads it. Where either is missing, the threshold is aspirational, and it is more dangerous than no threshold at all because it looks like a control. Setting appetite against what you can measure today, then scheduling the instrumentation work needed to raise it, is the honest sequence.

The frozen appetite. Set at deployment, never revisited, while the model drifts, the deployment scope creeps, and the regulatory floor rises underneath it. Test: when was appetite last changed, and what changed it? A statement never amended in an environment moving this fast is unattended.

What the board should approve

The board approves the dimensions, the consequence tiers, appetite per dimension, the Tier 4 positions specifically, the trigger actions, and the review cadence.

It should stay out of individual tolerance thresholds. Those belong to management, within the appetite the board has set. Pull that upward and you get a board deciding precision targets while nobody decides whether the organisation should be automating the decision at all.

Six months after this is signed off, a threshold will be breached. The pre-agreed trigger will fire, and nobody will need to convene a meeting to decide whether it mattered. Every part of the structure above exists to produce that one outcome, which is the same standard the wider risk register is held to.