Buyer guide

AI risk assessment: how to score a use case before launch

An AI risk assessment scores a single proposed use of AI before it is built, so the controls and the sign-off match what could actually go wrong. Score six factors from 1 to 3: the data the system reads, how much it can do without a person, its effect on people, the cost of an error, its exposure to untrusted input and outside users, and how often it runs. The total sets a risk tier, and a top score on autonomy or on effect on people makes it high risk whatever the total. Then list the specific failures, give each a control and an owner, and record a decision to proceed, redesign or stop. Where personal data is involved, the assessment feeds into the data protection impact assessment, which still has to be done.

AI risk assessment. A structured scoring of one proposed AI use case, done before it is built, that sets how much control it needs, who must approve it and whether it should proceed.

Checked . The ICO, NCSC, UK government and NIST sources this page cites were read on their publishers' own sites on 29 September 2026. The six-factor scale is 1AYM's own, not a published standard. The worked example is fictional: it illustrates the method and is not a client case, a result or a benchmark. This page is not legal advice.

What an AI risk assessment is for

An AI risk assessment answers one question about one use case: how much could go wrong, and so how much control does it need before it goes live? It sits between the idea and the build, and what it produces is a decision with conditions attached. A thirty-page report that nobody acts on is a sign the assessment was done for its own sake.

It is narrower than a governance framework. The framework is the organisation's rulebook: the policy, the register of AI uses, the risk tiers and the controls each tier requires. Our AI governance framework guide covers that, with a policy template you can copy. The assessment is how one use case gets placed on those tiers. If you have no framework yet, the tiers on this page will do as a starting point, and the framework can grow out of your first few assessments.

It also sits next to a data protection impact assessment (DPIA) rather than replacing it. Under UK GDPR, processing of personal data that is likely to result in a high risk to people needs a DPIA before the processing starts, and the ICO's guidance on AI and data protection says that in the vast majority of cases the use of AI will meet that bar. The ICO's list of processing likely to be high risk covers new technologies and the novel use of existing ones, naming AI, as a criterion that, combined with others, makes a DPIA necessary. The scores below give your data protection officer much of what a DPIA needs, but the DPIA stays theirs to run.

Do the assessment at the idea stage, well before the security review. A control designed in is a few lines of the build plan. The same control retrofitted after a reviewer blocks the launch means rework and a delay, and a late no from legal can cost the whole build.

Score the use case on six factors

Score each factor 1, 2 or 3 from the descriptions in the table, then add them up. The six are the things that change what a failure looks like: what the system can see, what it can do, who it touches, what a mistake costs, who can reach it and how often it runs. The scale is ours, so move the thresholds to fit your own risk appetite. Then keep them fixed, because the value is in comparing one use case with the next on the same terms.

Score the system as it will be built. The pitch is not the design. "A person reviews every output" only scores as a person reviewing every output if the design makes them do it, and if they will have the time.

The six factors of the AI risk assessment, with what scores 1, 2 and 3 on each
DimensionWhat you are askingScores 1Scores 2Scores 3
DataWhat is the most sensitive data the system can read or send out?Public information, or internal content with no personal dataPersonal data, or confidential business dataSpecial category or criminal offence data, children's data, financial account details, or data a contract restricts
AutonomyWhat can it do without a person deciding?Suggests or answers; a person decides what to do with itDrafts or prepares an action that a person reviews before it takes effectActs on a system, a record or a customer with no person in between
Effect on peopleCan its output change what happens to a person?No individual is affectedIt informs a decision a person makes about someoneIt decides, or in practice settles, someone's access to money, a service, a job, a benefit or their rights
Cost of an errorIf it gets something wrong, what does that cost, and can it be undone?Little, and caught before it mattersReal money, time or reputation, but reversibleCostly, hard to reverse, or a breach of a legal duty
ExposureWho can put input into it, and who sees what comes out?Internal staff only, working from trusted sourcesCustomers or suppliers see its output, or it reads documents from outsideIt reads untrusted content such as inbound email or web pages, and its output reaches customers or the public
ScaleHow often does it run?Up to about 100 times a weekMore than 100 a week, up to about 1,000 a dayMore than 1,000 a day, or with no cap

Exposure is the factor that gets left out. A system that reads inbound email reads text written by anyone who can send you an email, and the UK government's AI Playbook names prompt injection, alongside data poisoning, perturbation attacks and hallucinations, among the threats that are specific to AI. Untrusted input plus a system that can act is the combination to look for first.

From score to tier, and what each tier needs

Six factors scored 1 to 3 give a total between 6 and 18. Add one rule on top, a hard trigger: a 3 on autonomy or on effect on people puts the use case in the high tier whatever the total. Adding scores up is a convenience, and it hides exactly the case you most need to see, a small internal system that runs twenty times a week and decides something about a person.

Risk tiers from the total score, with what each tier needs before launch and who signs it off
DimensionScoreWhat it needs before launchWho signs it off
Low6 to 9, with no 3 on autonomy or effect on peopleAn entry in the AI register, an approved tool or supplier, access limited to the people who need it, and loggingThe business owner of the process
Medium10 to 13, with no 3 on autonomy or effect on peopleEverything in low, plus a written risk register, a test set of real past cases with pass thresholds, DPIA screening where personal data is involved, and a security review of what the system can read and changeThe business owner, with the data protection officer and security
High14 to 18, or any 3 on autonomy or effect on peopleEverything in medium, plus a full DPIA where personal data is involved, a deterministic check or a named approver on every consequential action, a switch that turns the capability off, a staged rollout and a review dateA named executive, after data protection, security, legal and risk have each signed their part

A high tier rarely means stop. It means the controls have to exist before the system does, and quite often the cheaper move is to redesign the use case until it scores lower, which is what happens in the worked example below.

For the actions that stay consequential after a redesign, the agent proposes, verifier gates pattern is how we build the check: the model proposes, deterministic rules decide whether the action goes ahead, and anything that fails goes to a person with the reason attached.

A worked example (fictional)

This example is invented to show the method. The retailer, the volumes and the design are illustrative, and none of them is a 1AYM client, a client result or a benchmark. A fictional online retailer receives about 3,000 customer service emails a week. The proposal is an AI assistant that reads each email, looks up the customer's orders, writes and sends the reply, and grants or refuses refunds of up to £75 on its own.

Scored as proposed, it lands in the high tier twice over: on its total, and on two hard triggers. So the team redesigns it before the build, which is cheaper than arguing about controls for a design nobody needed. In the redesign the assistant drafts every reply for an agent to send, can propose a refund but never grant or refuse one, and looks up orders only for the sender's own email address, a limit built into the lookup itself so the model cannot widen it.

Fictional worked example: the customer service assistant scored as proposed and as redesigned
DimensionAs proposedRedesignedWhy
Data22Names, addresses and order history are personal data. No special category data or card details: payments are handled elsewhere.
Autonomy32As proposed, it sends replies and moves money. Redesigned, an agent sends every reply and decides every refund.
Effect on people32Refusing a refund settles a customer's access to their money. Redesigned, the assistant informs a person's decision.
Cost of an error21A wrong refund or a wrong promise costs money, capped at £75 a time. Redesigned, an agent catches the error before it leaves.
Exposure33It reads inbound email from anyone, and its words reach customers. The redesign does not change who can write to it.
Scale22About 3,000 emails a week is roughly 430 a day.
Total1512The six scores added up, out of 18.
TierHighMediumAs proposed, high on the total and on two hard triggers. Redesigned, medium, with no hard trigger.

Notice what did the work. The biggest drop in risk came from the redesign, before anyone wrote a line of code, and none of it was exotic. That is the case for running the assessment while the idea is still cheap to change.

The risk register for the redesigned version

A medium score says how much control to apply. Choosing the controls takes a risk register: the specific ways this design could fail, rated for likelihood and impact before any control, then a control and an owner for each, then the rating once the control is in place. These are the four risks that matter most for the fictional assistant.

Fictional worked example: the four main risks, each with its rating before control, its control and owner, and its rating after control
DimensionBefore controlControl and ownerAfter control
Instructions hidden in a customer's email make the draft promise a refund or reveal another customer's orderLikely, high impactOrder lookups limited by the system to the sender's own account. Drafts checked for refund promises and for any order number that is not the sender's, with failures held for a supervisor. Owner: the engineering lead.Unlikely, medium impact
The draft states a returns policy that does not existLikely, medium impactReplies drawn from the current policy documents only. A test set of 300 past emails with known correct answers runs on every prompt or model change, with a pass threshold. Owner: the customer service manager.Unlikely, low impact
Agents approve drafts without reading themLikely, medium impactA weekly sample of sent replies reviewed, and the share of drafts sent unedited tracked, with a rising share treated as a warning rather than a success. Owner: the customer service team leader.Possible, low impact
The model supplier keeps or uses customers' personal data beyond what the contract allowsPossible, high impactSupplier terms on data use and retention checked before signing, a DPIA completed, and only the data a reply needs sent to the model. Owner: the data protection officer.Unlikely, medium impact

The decision: proceed to build at the medium tier, with the four controls as conditions of launch. Automatic refunds are not ruled out for good. They come back as a separate assessment once three months of logs show how often agents change the refund the assistant proposes, which is evidence the original proposal never had.

What legal, security and risk will ask before the build

A late block from legal, security or risk is usually a question nobody asked them early enough. The UK government's AI Playbook tells public servants to engage compliance, legal and data protection experts early, including during product development, and to work with commercial colleagues from the start. The advice travels well beyond government. Take each function the questions below, with the assessment, at the scoping stage.

What each function will ask before an AI build, and what to bring to answer it
DimensionWhat they will askWhat to bring
Data protection officerWhat personal data does it use, on what lawful basis, and is a DPIA needed? Where do automated steps affect people, and how much human involvement is there, at what stage? How long is data kept, and who else processes it?The data and effect-on-people scores, a data flow diagram, the supplier's terms on data use and retention, and a DPIA screening. The ICO expects the degree of human involvement in a decision to be identified and recorded, and a DPIA to start at the earliest stages of a project and be kept under review as a live document.
SecurityWhat can the system read and change, and with whose permissions? How could an attacker misuse it? What is logged, and who watches it? Which third-party models, data and code does it depend on?A threat model that covers AI-specific attacks, the list of tools and permissions it holds, the logging plan and the supplier list. The UK's Code of Practice for the Cyber Security of AI asks for threat modelling that addresses attacks such as data poisoning, model inversion and membership inference, an inventory of assets, a secure supply chain and monitoring of the system's behaviour. The NCSC's guidelines ask for restrictions on the actions an AI component can trigger.
LegalWhat does the supplier contract say about our data, our outputs and liability? Do we have to tell people they are dealing with AI? Does sector regulation apply? Could the output be used in the EU?The contract terms, the wording users will see, and the sector rules that apply. The EU AI Act can reach a UK organisation whose system's output is used in the EU, and the governance framework guide summarises its scope and timeline.
RiskDoes this fit our risk appetite? What risk is left after the controls, who owns it, and when do we look again?The completed template: the scores, the tier, the risk register with owners and ratings after control, the decision and the review date.
ProcurementHas the supplier been checked, and can we leave the contract and take our data with us?The supplier checks, the exit terms and the clause on returning or deleting data.

None of this is legal advice. It is the set of questions these teams usually ask, in an order that saves a rework. Your legal adviser and data protection officer decide what the law requires of you.

For US buyers: how the assessment maps to the NIST AI RMF

For US buyers the usual reference point is the NIST AI Risk Management Framework. AI RMF 1.0 was released on 26 January 2023 and is intended for voluntary use, and NIST says it is being revised as part of the White House AI Action Plan. It has four functions: Govern, Map, Measure and Manage. The template on this page covers the Map and Manage steps, and part of Measure, for a single use case; the governance framework is where Govern lives.

Map 1.1
Intended purposes, context and the laws and norms that apply are understood and documented. In the template, that is the description section.
Map 1.5
Organisational risk tolerances are determined and documented. That is the tier table and its thresholds.
Map 3.5
Processes for human oversight are defined, assessed and documented. That is the autonomy score and the approver named in each control.
Map 5.1
The likelihood and magnitude of each identified impact are identified and documented. That is the risk register.
Manage 1.1 to 1.3
A determination is made whether development or deployment should proceed; treatment is prioritised by impact, likelihood and available resources; and responses to high-priority risks can include mitigating, transferring, avoiding or accepting. That is the controls and the decision line.
Measure 2.6
The system is evaluated regularly for safety risks, its residual negative risk does not exceed the risk tolerance, and it can fail safely. That is the test set and the review date.

For generative AI, NIST's Generative AI Profile (NIST AI 600-1, 26 July 2024) names twelve risks that are unique to generative AI or made worse by it. The four risks in the worked example fall under four of them: information security, confabulation, human-AI configuration and data privacy. The list is worth reading when you write a risk register, because it includes risks that are easy to forget, such as harmful bias or homogenization, intellectual property, and value chain and component integration.

The framework is voluntary, and sector regulators, state laws and your own contracts may ask for more. Banking is one example: SR 26-2, the revised model risk guidance the Federal Reserve issued with the FDIC and the OCC, places generative and agentic AI outside its scope and leaves their controls to each bank's own risk management, and our guide to SR 26-2 and generative AI sets out a control set a model risk team can build on. This page is not legal advice on any of these.

For UK public bodies

Central government adds two expectations. The AI Playbook for the UK Government asks that humans validate any high-risk decisions influenced by AI and, where real-time review is not possible, as with a chatbot, that human control sits at other stages. And central government departments and arm's length bodies in scope are required to use the Algorithmic Transparency Recording Standard; the Playbook encourages other public sector bodies to use it too. If you are in scope, plan that record alongside the assessment.

This section summarises published guidance and is not legal advice. Our public sector page explains how UK public bodies can work with 1AYM, and what G-Cloud covers.

An AI risk assessment template you can copy

Fill this in at the idea stage, share it with the functions above at scoping, and update it before launch and at each review. A line you cannot fill in is a question to answer before the build.

AI risk assessment template
AI use case risk assessment: [use case name]

Business owner: [name, role]
Assessed by: [names, roles]
Date: [date]
Review by: [date, and after any change of data, tools, model, supplier, volume or users]

1. Description
What the system does, and for whom: [ ]
Data it reads, and where its output goes: [ ]
Actions it can take, and on which systems: [ ]
Supplier, model and hosting: [ ]
Volume: [ ] a week

2. Scores (1 to 3, with the reason for each)
Data: [ ]
Autonomy: [ ]
Effect on people: [ ]
Cost of an error: [ ]
Exposure: [ ]
Scale: [ ]
Total: [ ] of 18
Hard trigger (a 3 on autonomy or effect on people): yes / no
Tier: low / medium / high

3. Risk register (one line per risk)
Risk: [ ] | Before control: likelihood [ ], impact [ ] | Control: [ ] | Owner: [ ] | After control: [ ]

4. Sign-off questions
Data protection: DPIA needed? yes / no / screening attached. Human involvement recorded: [ ]
Security: threat model, tool permissions, logging plan, suppliers: [ ]
Legal: contract terms, what users are told, sector rules, use in the EU: [ ]
Risk: within appetite? yes / no. Owner of the remaining risk: [ ]
Procurement: supplier checks and exit terms: [ ]

5. Decision
Proceed / proceed with conditions / redesign / stop
Conditions of launch: [ ]
Signed: [name, role, date]

Where 1AYM fits

Scoring candidate use cases on value, feasibility and risk is part of 1AYM's AI Opportunity & Feasibility Sprint, which is fixed scope and typically takes two to four weeks. Its governance review establishes what has to be true for legal, security and risk to sign off before anything is built, so the questions in this guide get answered at scoping rather than in the last week.

The controls a medium or high score calls for are engineering: permissions the system enforces, checks on consequential actions, tests in CI and an audit trail. We build them as AI governance implementation, as part of our production AI systems work, and we take small fixed-scope statements of work as well as larger builds. On a government-accredited EdTech's speaking assessment, for example, the confidence gate is set to route about 15% of interviews to a human examiner, with a further 5% audited at random. If you have an assessment half done and want a second opinion on the scores, book the call below.

For engineers: the technical checks behind each score

The scores are only as honest as the design behind them. Before agreeing a tier, a technical reviewer should be able to confirm each item below from configuration, code or logs rather than from a slide.

Tool permissions
List every tool the system can call, the identity it runs as and the scope of each call. Enforce scope in the tool layer, not the prompt: a lookup keyed to the authenticated sender cannot be talked into returning another customer's order.
Untrusted input
Treat every field a third party can write (email bodies, attachments, web pages, documents in the retrieval index) as a possible prompt injection. Keep it out of the instruction channel where the model API allows, and never let it widen a tool's scope.
Output checks
Deterministic validation on anything that leaves the system: schema, allow-listed actions, amounts within limits, identifiers that belong to the requester. A failed check holds the item for review with the reason attached.
Evaluation set
Real past cases with known correct outcomes, adversarial ones included, run in CI on every change to a prompt, model, tool or retrieval source, with a pass threshold that blocks release.
Retrieval corpus
Record who can write to the documents the system retrieves from. Write access to the knowledge base is write access to the answers, which is how data poisoning reaches a system nobody retrained.
Logging
Per request: input reference, model and prompt version, retrieved sources, tool calls with arguments, check results, approver and outcome. Agree retention with the data protection officer, because the logs hold personal data too.
Fine-tuned models
If the system is fine-tuned on personal data, put model inversion and membership inference in the threat model, as the UK Code of Practice for the Cyber Security of AI asks.
Off switch
A flag that turns the capability off without a deploy, tested before launch, and a fallback route for the work while it is off.

Sources

  1. [1]ICO, When do we need to do a DPIA?, read 29 September 2026
  2. [2]ICO, Guidance on AI and data protection: what are the accountability and governance implications of AI? (last updated 15 March 2023; under review following the Data (Use and Access) Act), read 29 September 2026
  3. [3]ICO, AI and data protection risk toolkit, read 29 September 2026
  4. [4]NCSC, Guidelines for secure AI system development (27 November 2023), read 29 September 2026
  5. [5]NCSC, Guidelines for secure AI system development: secure design, read 29 September 2026
  6. [6]Department for Science, Innovation and Technology, Code of Practice for the Cyber Security of AI (31 January 2025), read 29 September 2026
  7. [7]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026
  8. [8]NIST, AI Risk Management Framework, read 29 September 2026
  9. [9]NIST AI Resource Center, the AI RMF Core, read 29 September 2026
  10. [10]NIST, AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (26 July 2024), read 29 September 2026

Frequently asked questions

What is an AI risk assessment?

A structured scoring of one proposed use of AI, done before it is built. It records what data the system uses, what it can do without a person, who it affects, what an error costs, who can reach it and how often it runs. From that it sets a risk tier, lists the main risks with a control and an owner for each, and ends in a decision to proceed, redesign or stop.

Is an AI risk assessment the same as a DPIA?

No. A DPIA is a legal requirement under UK GDPR for processing of personal data that is likely to result in a high risk to people, and it has to be done before the processing starts. An AI risk assessment also covers security, money and reputation, and applies when no personal data is involved. Where both are needed, run them together: the data, autonomy and effect-on-people scores feed the DPIA, and the DPIA stays with your data protection officer. This is not legal advice.

Who should carry out an AI risk assessment?

The business owner of the process, with someone who understands how the system will be built. Scoring autonomy and exposure honestly needs technical input; scoring effect on people and the cost of an error needs someone who knows the work. Data protection, security, legal and risk then review a finished draft.

How often should an AI risk assessment be reviewed?

Whenever something changes that could move a score: a new data source, a new tool or permission, a change of model or supplier, a rise in volume or a new group of users. Otherwise on a fixed date, and at least once a year. The ICO describes a DPIA for AI as a live document to review regularly, and the same reasoning applies here.

Do we need one for an off-the-shelf AI assistant?

Yes, but it is usually short. An assistant that only drafts text a person reads before using it scores low on autonomy, effect on people and the cost of an error. What pushes it up is data, meaning what staff paste in or connect it to, and exposure, meaning whether it reads outside content or acts in connected apps. Score it on what staff can actually connect it to.

Is there a standard AI risk assessment template?

There is no single mandatory template for private organisations in the UK. The ICO publishes an AI and data protection risk toolkit as a spreadsheet, and in the US the NIST AI Risk Management Framework is the usual voluntary reference. The template on this page is a short working version to adapt. Whatever you use, keep it the same across use cases so the scores can be compared.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM