Buyer guide

AI ROI: how to estimate it before you build and check it after

AI ROI is the value an AI system returns, minus everything it costs to build and run, divided by that cost over a stated period. Before you build, measure the work as it runs today (volume, time per case, error rate and cost per case) and estimate how much of it the system will change. Then count only the part that turns into money, because freed time is worth cash only when it removes spend, absorbs growth or earns revenue. Add the costs that tend to be left out, such as evaluation, monitoring and change management, and stress-test the result with worse assumptions. After launch, check it against the same baseline.

AI ROI. The net value an AI system returns over a stated period, realised benefit minus the full cost to build and run it, divided by that cost.

Checked . The UK government guidance this page cites was read on GOV.UK on 29 September 2026. The worked example is fictional: its figures illustrate the method and are not benchmarks, prices or results. This page is not financial, legal or tax advice.

The formulas: benefit, cost, payback and ROI

Five numbers decide whether an AI project pays back. The arithmetic is simple and the inputs are where cases go wrong, so write each input down as an assumption with its source next to it. The post-launch check later on this page tests those assumptions one by one.

Two of the inputs need defining. Loaded hourly cost is what an hour of the team's time really costs you: salary plus employer taxes, pension and overheads, divided by the hours people actually work. Realisation rate is the share of freed time that turns into money. Both come up again further down, and they are where optimistic cases usually go wrong.

The five AI ROI formulas, with what each one tells a sponsor
DimensionHow to work it outWhat it tells you
Gross annual valueHours freed a year × loaded hourly cost, plus error costs avoided, plus any revenue effectThe most the change could be worth if every freed hour became money.
Realised annual benefitGross annual value × realisation rateThe part that shows up in the accounts: spend removed, growth absorbed without hiring, revenue gained.
Net annual benefitRealised annual benefit minus annual running costIf this is zero or negative, the project never pays back, however cheap the build.
Payback periodOne-off cost ÷ (net annual benefit ÷ 12), in monthsHow long the money is at risk before the project has returned it.
ROI over a stated horizon(Total realised benefit minus total cost) ÷ total cost, over the same years, such as threeThe return per pound spent. Always state the horizon, because a three-year figure and a one-year figure for the same project are very different numbers.

Measure the baseline before anything is built

Every ROI figure compares the new way of working with the old one, so the baseline has to exist before the AI does. Measured afterwards, it becomes a reconstruction from memory, and the people who remember it best are the ones who want the project to have worked. I would not approve an AI business case with no measured baseline behind it.

The UK government's AI Playbook recommends early user research, with a business analyst looking for efficiencies in the current process, to establish a baseline of metrics that project outcomes can be measured against. HM Treasury's guidance on evaluating AI interventions goes further. Think about the baseline early, it says, and if the evaluation will compare the AI system with business as usual, document precisely what business as usual consists of before the system is introduced. It also warns that a baseline is harder to define when AI replaces complex or subjective work, which is an argument for starting early rather than for skipping it.

Measure the one workflow the project will change, in the units the ROI will be reported in. Four weeks of data is a sensible minimum for most volumes, because it is long enough to catch a month-end peak.

Volume
How many cases arrive a week or a month, from the system that logs them rather than an estimate.
Time per case
Average handling time, taken from system timestamps where they exist and checked against a short time sample where they do not.
Errors and rework
The share of cases that come back, get corrected or cause a complaint, and what each one costs to put right.
Cost per case
Time per case × loaded hourly cost, plus any per-case charges such as outsourced handling or penalties.
Cycle time
How long a case waits from arrival to done, if the value of the project is speed rather than effort.
Quality as the customer sees it
A satisfaction score or a complaint rate, so a faster process that is worse for customers shows up as a cost.

When hours saved become money

A saved hour is worth its loaded cost only if something changes because of it. If the same people are paid for the same hours and fill the freed time with other work, you have gained capacity rather than cash, and a business case that counts it as cash overstates the return.

So estimate a realisation rate: the share of freed time that will turn into money, and the route by which it will. The honest routes are a vacancy you no longer need to fill, overtime or agency cover you stop paying for, growth the team absorbs without hiring, errors that stop costing you refunds or penalties, and revenue from work that now gets done. Name the route for each part of the benefit and the person who will make it happen. A realisation rate with no owner tends to stay at zero.

HM Treasury's Green Book makes a related point about comparisons. Its business as usual option, the benchmark every proposal is compared against, does not mean doing nothing, because current arrangements carry their own costs, benefits and risks. For an AI case, compare with the process as it would otherwise run, including any hiring or growth you already planned.

The costs people forget

Everyone sees the build cost. The lines below are the ones that get left out, and several of them come back every year. The AI Playbook warns that, depending on the tool, AI implementation can have considerable compute costs, and that ongoing investment is needed to maintain these systems, update them and keep their performance stable over time. When a buyer writes requirements for suppliers, the Playbook asks for guidance on budget that considers hidden costs.

Costs to include in an AI business case, what each covers and whether it recurs
DimensionWhat it coversWhen it lands
Baseline and test setMeasuring the current process, and collecting real past cases with known answers to test the system against.Before the build
Evaluation rerunsRunning the test set again whenever a prompt, a model or a rule changes, and someone's time to read the results.Every change, for the life of the system
Monitoring and loggingTools and people to watch live performance, catch drift and keep a record of what the system did.Every month
Model usage and hostingPer-use model charges, compute and storage. Usage grows with volume, so price it at the volume you expect, not the pilot's.Every month
Change management and trainingRewriting procedures, training the people who use the output, and managers' time while the new way of working beds in.Mostly at launch, again with each big change
Security, data protection and legal reviewAccess design, a data protection impact assessment screening where personal data is involved, and contract review.Before launch, then when scope changes
Support and maintenanceFixing faults, updating integrations when the systems around it change, and moving to a new model version when the current one is retired.Every year
The process owner's timeThe person who knows the work, answering questions, supplying cases and checking results during the build and after it.During the build, then a few hours a month

Evaluation and monitoring are where the engineering that makes production AI systems safe to depend on shows up in a budget: regression tests, logs, failure handling and someone who reads them. Leave them out and the case looks better on paper while the system gets harder to trust in use. For work where a wrong answer costs money, the controls between the model and your systems belong in the cost line too. The pattern where the model proposes and a deterministic check decides is one way to build them.

Change management is the easy line to cut when a budget is squeezed, and it is the line that decides whether anyone uses the system. The AI Playbook asks organisations to plan for how AI will change the way people and processes work, including change management plans and training based on a learning needs analysis.

A worked example (fictional)

This example is invented to show the method. The business, the figures and the supplier quote are illustrative, and none of them is a 1AYM price, a client result or a benchmark. A mid-sized distributor's accounts payable team answers supplier invoice queries by email. The proposal is an AI step that reads each query, pulls the invoice and payment records, and drafts a reply for a person to check and send.

Assume a loaded hourly cost of £35 for the team and £50 for the process owner and manager. The coverage and handling-time figures come from a test on 300 past queries, which is the kind of evidence a case should rest on.

Fictional worked example: estimating the ROI of an AI step for supplier invoice queries
DimensionFigureHow it is worked out
Queries a month2,000Baseline, counted over four weeks in the email system
Handling time today12 minutesBaseline average, from timestamps checked against a two-week time sample
Queries the AI step can help with70%Test on 300 past queries
Handling time with the AI step6 minutesTest on the same queries, including the time to check the draft
Rework rate4% today, 2.5% expectedEach rework takes 30 minutes
Hours freed a month1552,000 × 70% × 6 minutes = 140 hours, plus 2,000 × 1.5% × 30 minutes = 15 hours
Gross annual value£65,100155 hours × £35 × 12 months
Realisation rate50%A vacancy left unfilled and less overtime, owned by the finance operations manager; the rest of the freed time goes to supplier reconciliations, not priced
Realised annual benefit£32,550£65,100 × 50%
One-off cost£36,800Build and integration £30,000 (a fictional supplier quote); baseline and 300-case test set £2,000; change and training £3,600; security and data protection review £1,200
Annual running cost£15,960Model usage £960 (2,000 queries a month at an assumed 4p each); hosting, logging and monitoring tools £2,400; monitoring and review, 6 hours a month at £50, £3,600; two rounds of re-testing after model or prompt changes £3,000; support and maintenance £6,000
Net annual benefit£16,590£32,550 minus £15,960
Payback periodAbout 27 months£36,800 ÷ (£16,590 ÷ 12) = 26.6 months
Three-year ROI15%Benefit £97,650 against cost £84,680 (£36,800 plus 3 × £15,960), a net £12,970

On these figures the project pays back inside three years, but only just, and the running costs eat almost half the realised benefit. Once the forgotten costs are in, expect that shape. It is a lot less exciting than the £65,100 a year of gross value a first slide would show. The next section tests whether the case survives worse assumptions.

Stress-test the case before you approve it

HM Treasury's Green Book, the UK government's guidance on appraisal, says appraisals should be explicitly adjusted for optimism bias, the proven tendency to be over-optimistic about key assumptions. It describes the pattern as costs that are typically higher, benefits that are typically lower and delivery that typically takes longer than first anticipated. It expects the size of the adjustment to come from an organisation's own record of past forecasts against what actually happened, so if you keep that record, use it.

The Green Book also asks for sensitivity analysis, which changes key assumptions to see how the result moves, and switching values, the value an assumption would have to reach for an option to stop being value for money. Both are cheap to run on a spreadsheet, and either one tells a sponsor more than a single ROI figure does.

Apply that to the fictional example. Cut the benefit by 20% and raise every cost by 20%, and the three-year result turns from a 15% return into a 23% loss. The switching value for the realisation rate is about 43%: below that, the project loses money over three years, whatever else goes right. The 20% figures are our illustration, not a Green Book value.

That tells the sponsor what the pilot is for. The test on past queries has already measured the model. The open question is whether half the freed time really becomes money, so that is what the pilot should measure, and the approval should depend on the answer.

Numbers that mislead

Four figures turn up in AI business cases and progress reports, and each says less than it appears to. Keep reporting them, but never let one of them stand in for the return.

Usage counts
Logins, prompts and active users measure activity. A system used a thousand times a day on work that did not need doing returns nothing.
Model accuracy on its own
The AI Playbook separates model metrics, which measure how well the technology performs, from service metrics, which show whether users' needs and business goals are met, and warns the two can be very different. A highly accurate model whose output people ignore or misread has no return.
Self-reported time saved
A survey answer is an estimate, made by people with no baseline in front of them. Use surveys to find where to look, then measure there.
Hours saved, counted as cash
Freed time becomes money only through a named route, such as spend removed or growth absorbed. Without the realisation rate, the figure measures capacity.

The post-launch check: did it pay back?

The check after launch reuses the worksheet. Measure what the baseline measured, the same way, and put the actual figure next to each assumption: volume, coverage, time per case, rework, realisation and each running cost from the invoices. The gaps show you which assumption was wrong, and that is worth more than a single yes or no.

Run it at about 90 days, once the new way of working has settled, and again at a year. HM Treasury's guidance on evaluating AI interventions recommends rapid evaluation in the early stages of a roll-out and regular evaluation after full roll-out, because AI systems keep changing after launch. Repeat the check after any large change of model or scope.

Where you can, compare against a group that has not got the system yet. The same guidance points out that because AI is delivered digitally, you can control who uses it and when, for example by giving one team access before another. Measuring both groups before and after, a method it names difference-in-differences, separates the system's effect from a quiet month or a change in the work itself. If the whole team switched over on the same day, all you have is a before-and-after comparison, and the report should say so.

End the check with a decision: continue, fix the assumption that failed, or stop. A project that misses its case and is stopped early still leaves the evidence for the next case.

For public bodies: the Green Book and spend controls

UK central government has a set method for this. The AI Playbook for the UK Government says any investment approaching £10 million will typically need a five-part business case, for which the Green Book must be used. Below that threshold, it says teams should strongly consider the Government Digital Service's guidance on agile business cases. It also says digital and technology spend above £100,000 for anything public facing, and above £1 million for anything else, must be assured through your assurance boards, following your organisation's own process.

The Green Book discounts future costs and benefits at the social time preference rate, set at 3.5% in real terms for years 1 to 30, and asks practitioners to keep monitoring and evaluation in mind throughout appraisal, with value for money evaluation among the types it names. HM Treasury's AI evaluation guidance adds that many AI interventions will need more substantial evaluation than similarly sized business as usual work, because they are untested and carry distinct risks.

This section summarises published guidance. It is not legal or financial advice, and your finance and commercial teams set the rules you follow. Our public sector page explains how UK public bodies can work with 1AYM, and what G-Cloud covers.

An AI ROI worksheet you can copy

Fill this in before approval, keep it with the business case, and complete the actual column at 90 days and at a year. Any line you cannot fill in before approval is a question for the pilot to answer.

AI ROI worksheet
AI ROI worksheet: [workflow name]

Sponsor: [name, role]
Process owner: [name, role]
Horizon: [for example, 3 years]

1. Baseline, measured before the build (source and dates for each)
Volume a month: [ ]
Time per case: [ ] minutes
Error or rework rate: [ ]%, cost per error: [ ]
Loaded hourly cost: [ ]

2. Assumptions (estimate / actual at 90 days / actual at 1 year)
Share of cases the AI step helps with: [ ] / [ ] / [ ]
Time per case with the AI step: [ ] / [ ] / [ ]
Error or rework rate with the AI step: [ ] / [ ] / [ ]
Realisation rate, with the route and its owner: [ ] / [ ] / [ ]
Revenue effect, if any: [ ] / [ ] / [ ]

3. One-off costs
Build and integration: [ ]
Baseline and test set: [ ]
Change management and training: [ ]
Security, data protection and legal review: [ ]

4. Annual running costs (estimate / actual)
Model usage and hosting: [ ] / [ ]
Monitoring and logging: [ ] / [ ]
Evaluation reruns: [ ] / [ ]
Support and maintenance: [ ] / [ ]
Process owner's time: [ ] / [ ]

5. Results
Realised annual benefit: [ ]
Net annual benefit: [ ]
Payback period: [ ] months
ROI over the horizon: [ ]%

6. Stress test
Result with benefits down [ ]% and costs up [ ]%: [ ]
Switching value of the realisation rate: [ ]%

7. Decision at each check: continue / fix [assumption] / stop. Reason: [ ]

Where 1AYM fits

Building this case is what 1AYM's AI Opportunity & Feasibility Sprint is for. It is fixed scope and typically takes two to four weeks: we look at how the work is done today and what state your data and systems are in, score candidate use cases on value, feasibility and risk, settle build versus buy, and write the roadmap, the investment case and the governance conditions. Whoever writes the case, make sure its assumptions are written down, because they are what make a post-launch check possible.

If a case needs evidence before full approval, a pilot built to measure coverage and realisation can be scoped on its own, because 1AYM takes small fixed-scope statements of work as well as larger builds. For a quick read of where you stand first, the free AI readiness assessment covers ownership, data, systems and controls in 13 questions. And if you already have a case and want a second opinion on its assumptions, book the call below.

For engineers: instrumenting the baseline and the ROI check

The finance model is only as good as the data behind it. Put everything below in place before the pilot starts, so the before and after figures come from the same instruments.

Case-level event log
One record per case with arrival, first touch, completion and rework timestamps, the handling path (AI-assisted or not) and an anonymised handler reference. Baseline and post-launch figures are then the same query over different dates.
Baseline from systems, not surveys
Derive time per case from ticket or workflow timestamps, and validate against a short time-and-motion sample where timestamps miss offline work. Record the sampling method with the figure.
Cost attribution per case
Tag model calls, compute and storage by workflow so the run cost per case can be read from billing data rather than apportioned by guesswork. Include evaluation runs, which are easy to lose in a shared account.
Holdout or staggered rollout
Gate the capability behind a feature flag by team or site, so one group can start later. That gives the comparison group a difference-in-differences estimate needs, and a clean way to switch the feature off.
Evaluation in CI
Keep the test set of real cases in the repository with its thresholds, and rerun it on every change to a prompt, model or rule. Record the cost of each run, because it is a recurring line in the case.
Quality guardrails
Track rework, escalation and complaint rates next to the speed metrics, so a faster process that sends more cases back shows up as a cost rather than a win.
Discounting
For multi-year cases, net present value is the sum over years t of (benefit minus running cost) ÷ (1 + r) to the power t, minus the one-off cost. Use your finance team's hurdle rate for r; UK government departments and arm's length bodies, which must follow the Green Book, use its 3.5% real rate for years 1 to 30. At 3.5%, the fictional example's three-year net present value is about £9,700.
Reconcile to finance
Agree the loaded hourly cost and the realisation evidence with the finance team before launch, so the post-launch figures are ones they will sign.

Sources

  1. [1]HM Treasury, The Green Book: UK government guidance on appraisal (2026 edition, last updated 5 February 2026), read 29 September 2026
  2. [2]HM Treasury Evaluation Task Force, Guidance on the Impact Evaluation of AI Interventions (Magenta Book supplementary guidance, updated 15 May 2026), read 29 September 2026
  3. [3]UK government, Artificial Intelligence Playbook for the UK Government (10 February 2025), read 29 September 2026

Frequently asked questions

How do you calculate AI ROI?

Take the realised benefit over a stated horizon, subtract the total cost of building and running the system over the same horizon, and divide by that total cost. Realised benefit is the value of freed time, avoided errors and extra revenue, reduced to the share that actually becomes money. Total cost includes evaluation, monitoring, change management and support, not only the build.

What is a good ROI for an AI project?

There is no reliable universal figure, so this page does not quote one. Judge a project against your own hurdle rate or cost of capital, against the other things the same budget could buy, and against how the result holds up when the key assumptions get worse. A modest return that survives a stress test is a better bet than a large one that needs every assumption to go right.

What goes in an AI business case?

The problem and the one workflow it affects, the measured baseline, the assumptions with their evidence, one-off and running costs including the ones that recur, the realised benefit with a named route to money and an owner, payback and ROI over a stated horizon, a stress test, the controls the system needs, and the plan for checking the result after launch. UK government departments and arm's length bodies also follow the Green Book, or below the typical £10 million threshold the agile business case guidance, and their own assurance process.

How soon after launch can you measure AI ROI?

Early signals such as coverage, time per case and rework can be read within weeks. Realisation, the part that turns freed time into money, usually takes longer, because it depends on decisions such as not refilling a vacancy. Check at about 90 days and again at a year, against the same baseline.

Can you measure the ROI of an AI assistant for all staff?

Only for specific tasks. A general assistant spreads small gains across many kinds of work, which is hard to measure as a whole. Pick a few named tasks where it is used heavily, baseline them, and give access to one group before another so you have a comparison. Treat the rest as capacity rather than a cash return.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM