August 23, 2026 · 7 min read
How to Measure AI ROI: Four Numbers to Capture Before You Build
By Sahan, co-founder, systems and delivery
AI ROI is decided before the build, not after it. Four numbers have to exist in writing before anyone writes a line of code: how often the process runs, what one run costs today in time and money, how often it comes out wrong, and what else in the business is changing at the same time. Capture those four and the return calculation afterwards is arithmetic. Skip them and you are arguing from memory against a baseline nobody wrote down.
That is the actual failure mode, and it is documented. Deloitte surveyed 1,854 senior executives at organisations already running AI implementations across Europe, the Middle East and the UK, published 22 October 2025. Only 15% of generative AI users reported significant, measurable return. 57% were using agentic AI, and only 10% reported significant ROI from it.
Deloitte’s explanation is not that the technology underperformed. Executives could not isolate what AI contributed, because it arrived at the same time as data quality work, team changes and process cleanup. That is a measurement problem, and it is fixable for free, before you spend anything.
The four numbers, in one place
Before a build starts, write down four things. Volume: how many times the process runs per week or month. Unit cost: the minutes one run takes and the loaded hourly rate of whoever does it. Error rate: how often a run comes out wrong and what fixing it costs. Attribution boundary: every other change landing in the same period.
| Number | How to capture it | What it is for | What breaks without it |
|---|---|---|---|
| Volume per period | Count from the system of record, never from an estimate | The multiplier on every saving | Savings sized off a guess, usually optimistic |
| Minutes and loaded cost per run | Time five real runs, apply the fully loaded rate | The numerator | Hours saved cannot be priced |
| Error and rework rate | Sample 50 completed runs, count the wrong ones | Stops you booking quality you never had | Quality regressions get recorded as savings |
| Attribution boundary | List every other change shipping the same quarter | Isolates what AI contributed | Someone else’s improvement gets credited to the agent |
Number one: how often the process actually runs
Count it from the system that records the work, not from what the team believes. Invoices issued, tickets closed, orders keyed, applications received. This count multiplies everything else, so an estimate that is 40% high makes the entire business case 40% high.
The number also tells you whether to build at all. A process that runs eleven times a month rarely repays a custom build, whatever the per-run saving looks like. The ranking method we use for that call is written up in what to automate first.
The UK’s Office for National Statistics found the same gap from the other side. In its release of 20 July 2026, covering businesses with 10 or more employees with fieldwork from 5 to 28 June 2026, the most common barrier to adopting AI was difficulty identifying business use cases, ahead of both cost and lack of expertise. Firms are not stuck on technology. They are stuck on naming a job they can measure.
Number two: what one run costs today
Time five real runs with a stopwatch, not one demonstration run by your fastest person. Then apply the fully loaded hourly rate, meaning salary plus employer costs plus the overhead the role carries, not base salary divided by 2,080.
Write it down with a date attached. “As of 12 August 2026, a shipping document set takes 22 minutes and is produced 140 times a month.” That sentence is worth more in six months than any dashboard you build in the meantime.
Number three: how often it comes out wrong
Pull 50 completed runs from the last quarter and count how many needed rework, carried a wrong figure, or went out and had to be corrected. That is your quality baseline.
Skip it and something predictable happens. The automation ships, the error rate stays flat or drifts slightly worse, and because nobody recorded the old rate, the whole thing gets booked as a win anyway. The quality baseline is the only defence against that, and it costs an afternoon.
Number four: what else is changing at the same time
This is the number that decides whether your final answer survives scrutiny. List every other initiative landing in the same window: the new system, the reorganised team, the process someone quietly tidied up, the two people who left.
Deloitte’s finding was exactly this. Executives could give a ballpark and nothing sharper, because the AI work shipped bundled with everything else. So either hold something back, or accept in advance that you are measuring the bundle and say so in the report.
Holding something back is easier than it sounds at small scale. You do not need a research design. Keep one branch, one shift, one product line or one queue on the old process for six weeks, then compare. Or stagger it: automate region A in March and region B in May, and region B’s March-to-May performance is your control.
Time saved is not money saved
The most common ROI shortcut is hours saved multiplied by hourly rate. The best available evidence says that conversion is not automatic.
The National Bureau of Economic Research published working paper 33795 in May 2025, revised November 2025, reporting a randomised field experiment across 66 firms and 7,137 knowledge workers. In the second half of the six-month trial, the 80% of treated workers who used the tool spent two fewer hours on email each week. The researchers did not detect a shift in the quantity or composition of the work those people did. Three of the four authors are at Microsoft Research and the tool studied was Microsoft’s, which is worth stating plainly.
Two hours a week is real. It becomes money only when the freed capacity is pointed at something that reaches the accounts: more orders handled without a hire, faster collections, a vacancy left unfilled. Pick which one before you build, then measure that.
The US Census Bureau’s Center for Economic Studies found the same pattern in firm data. Its April 2026 working paper, covering November 2025 to January 2026, reported that 66% of AI-using firms use it solely to augment tasks, and that AI-related employment decreases occurred in only 2% of firms. Headcount reduction is the exception, not the base case, so an ROI model resting on it is usually wrong before it starts.
The denominator most ROI advice never writes down
Return needs a cost side, and it has three parts: build cost, run cost, and human review cost. Run cost is published and checkable. Anthropic’s pricing documentation, accessed 23 August 2026, works a support example at roughly 3,700 tokens per conversation on Claude Haiku 4.5 and arrives at about $37.00 per 10,000 tickets, with batched work discounted 50% and cache reads at a tenth of the base input price.
Model cost is usually the smallest line on the page. Human review is usually the largest one people forget. If every output is checked by a person for the first eight weeks, that is a real cost and it belongs in the denominator.
One honest correction to all of the above: the first build pays for plumbing that later builds reuse, so judging build one on its own numbers understates the programme. That argument is set out separately in why your third automation costs half of your first.
What this looks like on our own builds
Every agent running in our own group had these numbers attached before it existed. That is why we can publish 41 registered agents, 9 live, and 5,450 runs at zero failures across 13 days, measured to July 2026, instead of a claim about efficiency. Each one is a tomte, our word for one production agent with one defined job, an owner and a kill switch, and each was scoped against a counted baseline rather than a hunch.
Capturing those four numbers is also most of what our AI operations audit does before it recommends building anything. $1,900, ten working days, and what comes back is a ranked build list with effort and payback estimates against measured volumes, not a strategy deck.
Send us the process you are considering automating, plus the four numbers if you have them. If you do not have them, say so, and we will tell you which to count first and how. Written reply within one business day, no meeting required. Start async.