July 24, 2026 · 7 min read
AI readiness assessment: why most are theater, and what works
By Anna, co-founder, build and content
An AI readiness assessment is supposed to tell you whether your business is prepared to adopt AI. Most of them do not. They hand you a maturity score, a colour-coded grid, and a slide deck, then leave. None of that gets a single agent into production. A useful assessment produces one thing: a ranked list of work worth automating, with effort and payback attached to each item, and a plan to build the first one.
We run this on our own three-company group before we run it for anyone else. The difference between the two kinds of assessment is not tone or thoroughness. It is whether you walk away with a score or a build list.
What an AI readiness assessment actually is
An AI readiness assessment reviews your data, tools, processes, and team, then judges how prepared you are to use AI. The standard version scores you across a few dimensions and places you in a maturity band: beginner, developing, advanced. The output is a report. The problem is that a report is an input to a decision, not the decision, and most of these never get anywhere near a working system.
There is a better version. It skips the band entirely and goes straight to the question you actually have, which is: what should we build, in what order, and what will it pay back.
Why most readiness assessments are theater
The failure data is now hard to ignore. In August 2025, MIT’s Project NANDA published “The State of AI in Business 2025” and found that 95% of enterprise generative-AI pilots delivered no measurable profit-and-loss impact. Only 5% captured real value. The report reviewed 300 public deployments alongside executive interviews and surveys, so this is not a small sample.
S&P Global Market Intelligence found the same rot from a different angle. In its 2025 “Voice of the Enterprise” survey of more than 1,000 respondents, 42% of companies said they had abandoned most of their AI initiatives that year, up from 17% in 2024. The average organisation scrapped 46% of its proof-of-concepts before they reached production. Gartner had predicted, back in July 2024, that at least 30% of generative-AI projects would be abandoned after proof-of-concept by the end of 2025. The measured number came in worse than the forecast.
Here is why this matters for assessments. Almost every company in those failure statistics was “ready” on paper. They had budget, executive support, and access to the same models everyone else has. A readiness score did not save them, because readiness was never the constraint. Picking the wrong work, and stopping at a pilot, was.
A maturity score measures the input, not the outcome
McKinsey’s “The State of AI in 2025”, published on 5 November 2025, reports that 88% of organisations now use AI in at least one function, up from 78% a year earlier. Adoption is nearly universal. Yet only about 6% of respondents qualify as high performers who attribute more than 5% of company earnings to AI. High adoption, thin results. That gap is the whole story.
The same survey names the thing that separates the two groups. High performers are 2.8 times more likely to have fundamentally redesigned workflows, 55% of them versus 20% of everyone else. Not a higher maturity score. A change to how the work moves. An assessment that grades your data hygiene and your cloud setup is measuring the wrong variable. The variable that predicts value is whether you are willing to rebuild a process around the agent, instead of bolting the agent onto the process you already have.
Focus beats breadth, and the data says so
Boston Consulting Group’s “Closing the AI Impact Gap”, published on 30 September 2025, surveyed 1,250 executives and found that only 5% of companies generate value at scale, while 60% see minimal gains. The leaders did something specific and countable: they prioritised an average of 3.5 use cases, against 6.1 for everyone else, and expected roughly 2.1 times the return. BCG also attributes about 70% of AI value to people and process change, and only 10% to the algorithms themselves.
Fewer bets, deeper build, more return. That is the opposite of what a broad readiness assessment encourages, which is a long horizontal survey of everything you could theoretically do. A ranked shortlist of three or four jobs, built properly, is worth more than a strategy deck listing twenty.
What a real AI assessment should hand you
Skip the maturity band. The deliverable that changes anything is a build list. When we run an AI operations audit on our own companies or a client’s, the output is written and specific: a map of how work actually flows today, a ranked list of automatable jobs with an effort, cost, and payback estimate on each, and a 90-day plan that says what to build first, what to skip, and what to buy off the shelf instead of building at all.
Here is the difference laid out plainly.
| Theater assessment | Build-list audit |
|---|---|
| Maturity score and a colour grid | Ranked list of jobs to automate |
| ”You are at level 2 of 5" | "Automate lead triage first, ~5 days, pays back in month one” |
| Twenty possible use cases | Three or four scoped, with effort and payback |
| Slide deck, then silence | A 90-day plan and a fixed quote for anything on it |
| No decision on what NOT to do | Explicit “skip this” and “buy this off the shelf” calls |
| You still do not know where to start | You start Monday |
We wrote a separate piece on the exact ranking method we use, scoring each process by effort against payback so the cheap, boring, high-frequency work rises to the top. That is usually where the first real return hides, not in the flashy customer-facing project.
The number that predicts whether you land in the 5%
If you take one thing from the 2025 research, take this: the companies that got value did fewer things and redesigned the workflow around each one. McKinsey’s 55%-versus-20% workflow-redesign gap and BCG’s 3.5-versus-6.1 use-case gap are the same finding from two firms. An assessment worth paying for tells you which three or four jobs to redesign, in what order, and what each is worth. It does not tell you that you are “developing” on a five-point scale.
Reliability is the other half of landing in the 5%, because a pilot that works once and fails in week three counts as one of the 95%. Our own group runs a registered fleet of 41 agents. In one measured 13-day stretch they ran 5,450 times with zero failures. We keep the full record of how that fleet is governed: every agent has an owner, a fence it operates inside, and a run log, and failures page a human instead of hiding. An honest assessment should tell you not just what to build, but what it takes to keep it running after the launch demo. Most skip that entirely, which is one reason so many pilots die.
How to tell a real assessment from a costume
Ask three questions before you pay for any AI readiness assessment. First: does the deliverable include a ranked build list with effort and payback per item, or does it stop at a score? Second: does it tell me what NOT to build, including what to buy off the shelf, or does it only sell more work? Third: is there a fixed price and a guarantee attached, or is it billed by the hour with the real proposal waiting at the end?
Our audit is $1,900, fixed, delivered in ten working days, with a full refund if we find nothing worth automating. That has not happened yet, but the guarantee stands, because an assessment that cannot justify its own price on the strength of its build list is theater with an invoice attached. You own everything it produces, whether you build with us afterwards or take the plan in-house.
Start with the work, not the score
You do not have a readiness problem. You have a “which three jobs, in what order, and what will they pay back” problem, and a maturity grid answers none of it. The 5% who got value in 2025 picked a short list and rebuilt the workflow around each item. That is the whole method.
If you want the ranked build list instead of the colour-coded slide, describe what is eating your team’s time at /start. You will get a written reply within one business day. No call, no discovery meeting, just the plan.