Set Your AI Risk Appetite Before You Approve Another Tool

AI Trust & Governance

6 October 2026 | By Ashley Marshall

Quick Answer: Set Your AI Risk Appetite Before You Approve Another Tool

UK organisations should define AI risk appetite before selecting controls or approving tools. Set measurable thresholds by risk type, name the decision owner, and use evidence from pilots to decide whether to accept, reduce, transfer or avoid each risk.

Most AI governance gets stuck because nobody has agreed which risks the business will accept. A measurable risk appetite turns vague caution into faster, more consistent decisions.

The missing decision is not whether AI is risky

Most leadership teams already accept that AI creates risk. The unresolved question is how much of each risk the organisation is prepared to carry in exchange for a worthwhile benefit. Without that answer, every proposal becomes a fresh argument. Security asks for more controls, operations asks for speed, legal asks for evidence, and the project sponsor cannot tell which concerns are decisive.

The UK government's AI Risk Management Toolkit, published on 8 September 2026, puts risk appetite before treatment and mitigation. It says appetite should be set at organisational level where possible, aligned with wider risk appetite and signed off at board level. That sequence matters. A control is not automatically good governance. It is useful only if it reduces a material risk to a level the organisation has agreed it can tolerate.

In practice, leaders should separate at least seven dimensions: financial loss, legal and regulatory exposure, contestability and redress, transparency, fairness, technical robustness, and security. The acceptable position will differ by use case. A drafting assistant used by a trained internal team can tolerate occasional low-impact errors if every output is reviewed. A system ranking job applicants, changing insurance pricing or prioritising patients needs a much lower tolerance for unfairness, opacity and weak routes of appeal.

This is not a demand for a perfect forecast. It is a demand for an explicit decision. Write down what outcome would be unacceptable, who can approve an exception, what evidence is required and when the position will be reviewed. That creates a usable boundary for teams instead of a policy statement that merely says to use AI responsibly.

Translate appetite into thresholds people can actually use

Words such as low, cautious and responsible sound sensible in a board paper but do not help a product owner decide whether a pilot can go live. A usable appetite statement links a risk category to a threshold, an evidence source and an authority. For example: no customer-facing answer may trigger a financial transaction without confirmation; fewer than one in 1,000 reviewed outputs may contain restricted personal data; every adverse employment recommendation must receive documented human review before it affects a candidate.

The government toolkit uses a five-point scale for both likelihood and impact. A likelihood score of 1 represents a probability below 5 per cent, 2 covers 5 to 20 per cent, 3 covers 20 to 50 per cent, 4 covers 50 to 80 per cent, and 5 is above 80 per cent. Impact also runs from negligible to catastrophic. Multiplying the two scores creates a common starting point, but it should not become fake precision. A score of 10 can describe very different situations, so the underlying scenario and affected people must remain visible.

A practical register should record the failure scenario, the population exposed, estimated frequency, maximum credible impact, current controls, evidence quality, owner, treatment decision and review trigger. Add a confidence rating beside the likelihood estimate. A pilot based on 200 test cases should not be treated as if it carries the same certainty as twelve months of production monitoring.

Set hard limits where rights, safety or regulatory duties require them. Use softer tolerances where trade-offs are legitimate. The point is not to compress every judgement into one red, amber or green cell. It is to make clear which number prompts additional testing, which result requires executive acceptance, and which outcome stops deployment. That is what turns risk appetite from governance language into an operating control.

Use evidence from the workflow, not confidence from the vendor

AI suppliers can provide model cards, security certificates, benchmark results and contractual promises. Those are useful inputs, but they do not establish the risk of your workflow. The same model may be low risk when it summarises public documents and high risk when it drafts a benefits decision from sensitive records. Your exposure depends on the data, prompt, integrations, user behaviour, review process and consequence of a wrong result.

The toolkit recommends combining historical data, model analysis, expert judgement, experimentation and monitoring. That gives UK businesses a sensible evidence ladder. Start with past complaint volumes, manual error rates, fraud losses and service downtime. Run a representative test set against the proposed workflow. Include edge cases and deliberately difficult examples. Ask domain experts to judge errors by consequence, not just accuracy. Then monitor the live service against the same thresholds.

The Department for Work and Pensions offers a concrete benchmark in its AI Security Policy, updated on 7 August 2026. It requires a Data Protection Impact Assessment where an AI tool processes personal data, enters an existing personal-data process, or produces output that affects people. It also requires users to check the accuracy, reliability and credibility of AI output. Suppliers must declare which tools are used and provide evidence of how output was achieved.

What this means in practice is simple: put evidence requirements into the approval gate. Do not accept 'the vendor says it is safe' as a test result. Ask for a workflow evaluation, a record of known failure modes, proof that escalation works, and a named person who will watch the production measures. If those artefacts do not exist, the risk is not yet understood well enough to approve.

Connect risk appetite to rights and regulatory expectations

Risk appetite does not allow a board to waive the law. A business cannot decide that discrimination, unlawful processing or the absence of required safeguards is acceptable because the commercial upside looks attractive. Appetite operates inside legal and contractual boundaries, then guides choices where judgement remains.

The Information Commissioner's Office recently published its response to its generative AI consultation series. Across five areas, it received 192 organisational responses and 22 responses from members of the public. The ICO retained its positions on purpose limitation, accuracy and controllership, and said developers relying on legitimate interests for web scraping may struggle with the balancing test where processing is invisible and mitigation evidence is weak. For deployers, the wider lesson is that contracts and vendor terms do not replace demonstrable compliance.

The ICO is also moving towards a statutory code on AI and automated decision-making. An August 2026 legal analysis from Arnold and Porter notes that the regulations requiring the code came into force on 12 May 2026. It also highlights that existing UK GDPR penalties can reach the higher of £17.5 million or 4 per cent of global annual turnover.

What this means in practice is that the risk register needs a separate compliance gate. Record lawful basis, purpose, data categories, affected groups, transparency information, human involvement, contestability and redress. If a proposal cannot meet a mandatory requirement, mark it as blocked rather than assigning a higher risk score. This distinction keeps commercial trade-offs honest and stops a colourful scoring matrix from disguising a legal gap.

Choose one of four treatments and fund it properly

Once a risk is described and compared with appetite, the owner needs a treatment decision. The government toolkit uses four familiar options: avoid the exposure, reduce it through controls, transfer part of it, or accept it. Teams often default to reduction because adding a filter, reviewer or approval step feels constructive. That can be the wrong economic decision.

Avoidance is appropriate when the potential harm is severe and mitigation is unreliable or disproportionately expensive. A small employer may decide not to use automated candidate ranking because it cannot validate outcomes across relevant groups. Reduction might mean limiting an assistant to retrieval from approved sources, removing tool permissions, adding human confirmation, or restricting the user population. Transfer can include contractual indemnities, insurance and service-level commitments, although accountability and reputation rarely move completely to a supplier. Acceptance is legitimate when the residual exposure sits inside the agreed threshold and the benefit justifies it.

Each treatment needs a cost, an owner and a deadline. A control without operational funding will decay. Human review requires capacity planning and sampling standards. Monitoring requires logs, dashboards and someone authorised to respond. Vendor commitments need contract management. Incident response needs a tested route to pause the system, preserve evidence and notify the right people.

There is also a valuable counterargument: formal risk scoring can slow useful experiments and reward teams that produce paperwork. That happens when the process applies the same burden to every use case. The answer is tiering. Give low-impact, reversible internal experiments a short form and rapid approval. Require deeper evidence when an AI system affects people, moves money, changes records, accesses sensitive data or acts through business tools. Proportionate governance is not lighter thinking. It is matching assurance effort to the maximum credible harm.

Build a 30-day risk appetite sprint

A mid-sized organisation does not need a six-month governance programme before it can make better decisions. It can establish a credible first version in 30 days. In week one, inventory the ten to twenty AI uses that matter most, including unofficial tools, supplier features and automations embedded in existing software. Group them by consequence rather than vendor.

In week two, hold a working session with an executive sponsor, operations, security, data protection, legal or compliance, finance and two frontline users. Define unacceptable outcomes and tolerable ranges across the main risk categories. Use real scenarios: an incorrect internal summary, a leaked customer record, an unfair candidate rejection, a payment sent to the wrong supplier, or a customer unable to challenge a decision. Name the person who can accept residual risk at each level.

In week three, score three representative workflows and test the thresholds. If every proposal is blocked, the appetite is probably too vague or the controls are immature. If every proposal passes, the thresholds are probably meaningless. Adjust the language until two competent reviewers can reach broadly the same conclusion using the same evidence.

In week four, connect the decisions to procurement, change control and live monitoring. Require a fresh review when the model changes, the data source expands, a new tool permission is added, the affected population changes or an incident exposes an assumption. Publish a one-page decision guide so staff know which uses can proceed, which need review and which are prohibited.

Finally, treat the first version as a controlled baseline, not a finished policy. The ICO's July 2026 discussion of regulatory sandboxes makes the useful point that experimentation works only when organisations bring hard problems and accept limits where privacy, safety and rights are at stake. Your internal governance should do the same: enable evidence-led testing, while making the stopping conditions unmistakable.

Frequently Asked Questions

What is AI risk appetite?

AI risk appetite is the type and amount of AI-related risk an organisation is willing to accept in pursuit of its objectives. It should be expressed through usable thresholds, owners and evidence requirements rather than broad statements about being cautious or innovative.

Is AI risk appetite the same as an AI policy?

No. A policy sets rules and responsibilities. Risk appetite explains how much exposure is tolerable and who can approve exceptions. The appetite should shape the policy, approval process and controls.

Who should approve AI risk appetite?

A board or appropriate executive committee should approve the organisation-level position. Named operational owners can then accept residual risk within delegated limits, with higher-impact decisions escalated.

How often should AI risk appetite be reviewed?

Review it at least annually and whenever a major incident, regulation, model capability or business strategy changes. Individual workflows also need review when their data, permissions, model or affected population changes.

Can a business accept a risk that breaches data protection law?

No. Risk appetite operates within legal and regulatory requirements. A mandatory safeguard or lawful processing requirement is a gate, not a commercial risk that can simply be accepted.

Do small businesses need a five-point scoring model?

Not necessarily. A smaller firm can use three levels if the definitions are clear. The important points are consistency, evidence, named ownership and a clear difference between acceptable, escalated and prohibited uses.

What evidence should support an AI risk score?

Use representative workflow tests, historical process data, complaint and incident records, expert judgement, vendor assurance, security testing and live monitoring. Record the confidence level when evidence is limited.

Does human review always reduce AI risk?

Only when the reviewer has time, competence, authority and enough information to challenge the output. A token sign-off can add delay without providing meaningful protection.