Back

AI Maturity Assessment: What the Score Cannot Tell You

A maturity level describes where your organisation is today. It does not tell you what it will do next, and the difference decides your programme.

Date

Reading time

8

min

Amelia Miller

Co-founder and CEO

An AI maturity assessment scores your organisation against a staged model of AI capability, and the score it hands back describes where you are without predicting what you will do next.

Those two things come apart most sharply in the organisations keenest to have the number. A maturity level is a description of a state. Adoption is a behaviour. No five-level model explains how a company moves from one level to the next, and that is the only question the person signing off the assessment actually has.

The number rarely stays a number, either. It becomes a slide, then a target, then a programme. If you are a transformation lead who has just been walked through a staged model by a consultancy, you are not really deciding whether the model is accurate. You are deciding whether to buy the engagement attached to it.

What a maturity score is good at

A maturity score is good at giving a leadership team one vocabulary for a problem they have been describing five different ways.

That is not a small thing. Before the score exists, the COO says the data is a mess, the CTO says the tooling is fine but nobody uses it, and the head of L&D says people are scared. All three are describing the same organisation and none of them can tell whether they disagree. A shared scale turns that into one conversation.

It also makes progress legible to a board. Boards ask for movement, and "we went from 2.1 to 2.8" is an answer, where "the pilots are going quite well" is not. And a stable scale lets you compare this year with last year, which matters if the programme runs longer than the attention span of the people funding it.

Keep all of that. The argument here is not that maturity models are worthless. It is that they are being asked to do a second job they were never built for.

A level describes you, it does not cause what happens next

Maturity levels are descriptive categories rather than causal mechanisms, which is why moving up one is not a plan and cannot be executed.

MITRE's AI maturity model runs through five levels, from Initial to Adopted, Defined, Managed and Optimized. Those are real distinctions and the model is a serious piece of work. But notice what the levels are: they are labels for clusters of characteristics that tend to occur together in organisations that have already got somewhere. They describe the destination, in the past tense, for companies that arrived by some route the model does not record.

So when a report says you are at level 2 and the roadmap says get to level 3, what has been handed to you is a photograph of somebody else and an instruction to look like it. Nobody can action "be more Defined". The work that actually moves a company is specific, local and usually unglamorous: one data pipeline nobody owns, one procurement rule that blocks the tool, one manager who has not told their team the tool is allowed.

This is the difference between a score and a diagnosis, and it is worth being precise about it, because ivee sells a readiness diagnostic and has an interest here. A diagnostic that names the one thing blocking you is a different object from a score that ranks you against a curve, and the four things a readiness assessment actually measures are chosen to be fixable. The first is executable. The second is a position.

Why the score climbs every year without anything changing

Most maturity assessments run on self-reported inputs, and self-reported inputs drift upward.

The mechanism is not dishonesty. It is that the people answering the questionnaire in year two have learned what the questions mean. They know what "documented governance process" is supposed to look like, because the year-one report told them. They answer more generously, more fluently, and with more of the vendor's vocabulary. The organisation has not changed. Its command of the language has.

There is a second, quieter effect. The person completing the assessment is very often the person who owns the AI programme, and therefore the person whose year is judged by whether the number went up. That is not a plot. It is an ordinary incentive acting on an ordinary human being holding a form with no external check on it.

Which is why a rising maturity score sits comfortably alongside a stalled programme. Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, naming unclear business value and poor data quality among the causes. None of those abandonment conditions would necessarily show up as a lower maturity level. They show up as projects dying.

The average hides the one thing blocking you

A composite score is an average, and an average is designed to smooth out exactly the outlier that decides your outcome.

Take a company scoring 4 out of 5 on strategy, 4 on leadership sponsorship, 4 on tooling, 4 on culture, and 1 on data access, because the customer records live in a system two people understand and neither will grant write access. The composite reads 3.4, which sounds mid-table and unremarkable. The reality is that nothing ships, because the one dimension scoring 1 is a gate and the other four are not.

Gates do not average. A constraint that blocks everything downstream cannot be compensated for by strength somewhere else, and the arithmetic of a composite score assumes precisely that it can. The organisation scoring 3.4 with a blocked gate and the organisation scoring 3.4 with four mediocre dimensions and no gate are in completely different positions, and the number cannot tell them apart.

This is also where the evidence problem bites. In ivee's own research, 4% of AI budget holders can prove their return on their organisation's AI spend, from a survey of 500 UK AI decision makers conducted in August 2026. The other 96% cannot show the numbers. A maturity score is frequently recruited to fill that hole, because it is the only figure in the room. It is not evidence of return. It is a description of shape.

Measure the constraint, not the score

The alternative is to find the single thing blocking you, size it, fix that one thing, and measure the thing you fixed.

In practice that is four steps, and none of them needs a five-level model:

  1. Name the binding constraint. Ask the people who have tried to use the tools what stopped them, and keep asking until the same obstacle comes back from three different directions. It is usually access, permission or ownership rather than skill.

  2. Size it. How many people does it block, how often, and what were they trying to do. A constraint blocking four people once a quarter is not the same problem as one blocking forty people daily, and the maturity score treats them identically.

  3. Fix that one thing. Not the roadmap. The one thing. Most of these are a permissions change, a named owner, or a manager saying out loud that the tool is sanctioned.

  4. Measure the constraint, before and after. If the obstacle was access, the measure is how many people have access. If it was hours lost to a manual step, the measure is hours.

That last step is where most programmes slide back to the score, because a real before-and-after needs a baseline and nobody took one. If the constraint you land on is time, ivee's calculator that turns hours saved per person into an annual figure will give you a defensible number in about a minute, which is a better input to a board conversation than a decimal place on a maturity scale.

None of this requires you to throw the framework away. If you already have a maturity report, read it as a checklist of things somebody has looked at, find the lowest score, and treat that as the candidate constraint. You are using the model as a diagnostic instrument rather than a scoreboard, which is what it is good for.

Where this argument does not hold

There are three situations where a stated maturity position is the right thing to have, and this argument gets weaker in all of them.

Where a standard or a regulator requires a documented position. ISO/IEC 42001, the international standard for AI management systems, requires an organisation to establish, implement, maintain and continually improve a documented management system, and certification against it is audited by an accredited body rather than self-declared. If you are going through that audit, a structured self-assessment against a defined scale is not a vanity metric. It is the artefact.

In very large organisations. Across thirty business units on four continents, a shared scale may be the only practical coordination device you have. Nobody can hold thirty binding constraints in their head. A common scale is a compression format, and at that size compression is worth the loss of fidelity.

Where the model is used as a checklist rather than a scoreboard. A framework that prompts you to look at data access, procurement, permissions and training in turn is doing useful work, and the six frameworks worth comparing mostly cover the right territory. The failure is in the aggregation, not the questions. Strip the score and keep the question set and you have lost nothing that mattered.

The argument holds where a mid-sized organisation is being sold an ascent up a scale as though the ascent were the work. That is the common case, and it is the one worth resisting.

If you already suspect you know which single thing is blocking you and want it named, sized and tested rather than scored, talk to ivee about a readiness diagnostic. Bring the maturity report if you have one. The lowest number in it is usually the right place to start.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.