Measuring AI adoption means turning what your team does with AI into a number that holds up when somebody outside the programme asks where it came from.
Three approaches are worth your time, and each answers a different question. Seat and usage activity tells you whether people opened the tool. Self-reported hours saved tells you how your team feels about it. An assessed capability measure tells you what they can now do. Choose by the question you have to answer, not by what is cheapest to collect.
Most teams do the opposite. They pick the measure that arrives on its own, present it, and find out in the meeting that they cannot defend it. The measure that gets collected and the measure that survives a question from finance are rarely the same measure, and the distance between them is where AI programmes lose their next round of budget.
The three measures at a glance
Criterion | Seat and usage activity | Self-reported hours saved | Assessed capability |
|---|---|---|---|
What it collects | Licences assigned, active users, prompts submitted, per tool and per department | Estimates from staff, usually one survey question on hours saved per week | Scored performance on a task the role actually contains, marked against a written rubric |
Collection cost | Near zero. It is already sitting in your admin centre | Low. One survey, about 10 minutes per person | High. Somebody designs the task, somebody marks it, and it runs twice |
How easily it is gamed | Easily, and usually by accident. One prompt in a month counts as active | Easily. The number is a feeling, and the feeling runs optimistic | Hard. Gaming a marked task means learning to do the task |
What it proves to a sceptical finance director | That people opened the tool. Nothing about what came out of it | Very little, once they ask how you know | That capability moved, on evidence they can re-run themselves |
Time to first signal | Days. Usage data lands within 48 hours | One survey cycle, so about a week | Six to 12 weeks, because you need a before and an after |
What does seat and usage activity actually measure?
Seat and usage activity measures whether somebody opened the tool, and nothing past that. It is a habit measure wearing a value measure's clothes.
Read the definitions before you quote the numbers. Microsoft's documentation for the Microsoft 365 Copilot admin reports counts an active user as anybody who submitted at least one prompt during the selected period, across a 7, 28, 90 or 180-day window. So one prompt in 28 days makes a person active, and your adoption rate is partly a function of which window you picked. ChatGPT Enterprise, Claude Enterprise and Gemini in Google Workspace all publish comparable admin-side counts with comparable caveats.
Licences assigned is not adoption at all. That is a procurement number, and reporting it as adoption is how a team ends up defending a figure that only ever described a purchase order.
What activity data is good for is finding the holes. Break it down by department and you will see which teams never started, which single team is carrying the average, and which tool people abandoned in week three. That is a real and useful finding, available for free, and no other measure gives it to you. Just do not put it on the slide where the question is whether the money worked.
What does self-reported hours saved actually measure?
Self-reported hours saved measures what your team believes about AI, which is a different quantity from what AI did. It is the number most often presented and the weakest number in the room.
The evidence on this is unusually direct. In a randomised controlled trial published by METR in July 2025, 16 experienced open-source developers worked through 246 real tasks from their own repositories. Beforehand they forecast that AI tools would make them 24% faster. Afterwards, having done the work, they estimated it had made them 20% faster. Measured, they were 19% slower. The direction of the effect was wrong, not just the size of it.
There is a second reason the number inflates. Workday's global study Beyond Productivity, fielded by Hanover Research across 3,200 leaders and employees in November 2025, found roughly 37% of the time AI saves goes straight back into correcting, verifying and rewriting its output. People report the saving. They rarely report the rework, because the rework does not feel like the same task.
None of which makes the survey worthless. It is the cheapest way to find out which specific tasks your team thinks are improving, and the free-text box is the best friction detector you will get for £0. Use it to decide where to look. Do not use it as the headline.
What does an assessed capability measure prove?
An assessed capability measure proves that a named person can now do a specific piece of work they could not do as well before, on evidence a sceptic can re-run. It is the only one of the three that answers the question a board is actually asking.
The mechanics are ordinary. Take a task from the job itself: a research analyst turning three anonymised filings into a one-page brief, a support lead drafting replies to 10 historic tickets, an ops manager reconciling a messy spreadsheet. Write a rubric with four or five criteria and a scale. Have somebody who knows the work mark it. Run it before the training and again after, on different material of the same type.
Be honest about the cost, because it is the reason most teams skip it. Designing one role-specific task and its rubric is roughly a day of a competent person's time. Marking runs 10 to 20 minutes per person per round, and there are two rounds. For 30 people that is about three days of effort in total, spread across a quarter. That is not nothing, and it is considerably less than the cost of a training programme nobody can defend.
It has a real limitation and it matters: capability is not use. Somebody can score beautifully on a marked task in July and never touch the tool in August. That is exactly why the activity data from the first measure belongs alongside it, doing the job it is good at. One measure tells you they can. The other tells you they do.
Which of the three survives a question from your finance director?
Only the assessed capability measure survives, and only if the baseline was taken before anybody was trained.
The question that kills a measurement is never "what is the number". It is "how do you know". Run each one through it. Usage activity answers "people opened it", which invites the follow-up you do not want. Hours saved answers "the team told us", and the moment somebody in the room has read the METR result, that is the end of the conversation. Assessed capability answers "here is the marked work from June, here is the rubric, here is the same task in September, marked by the same partner". A finance director can pull the folder and check.
One thing no measure here fixes, and you should say so before anyone else does: none of the three isolates AI from everything else that changed that quarter. You hired two people, a system got replaced, a big client left. A held-constant task gets closest, because the work is the same even when the context is not, but it is a comparison rather than a proof. Saying that out loud costs you nothing and buys you the rest of the number.
What does a usable baseline look like?
A usable baseline is the same measure, taken on the same people, before anybody is trained. Everything else is a reconstruction.
This is where most programmes come apart. The baseline gets taken after the training day, because that is when somebody remembers to ask, and a baseline collected after the intervention is just a second reading with a story attached. Two weeks before the first session, capture three things: the previous 28 days of usage activity from your admin centre, the scored task, and the one-question survey. Then repeat all three at 12 weeks. If you are running a six-week pilot with two departments, the baseline covers both departments and a third that is not being trained yet, which gives you something closer to a control group for free.
Here is the shape of a measure that survives challenge. The figures below are illustrative, not data, and yours will look nothing like them.
A 60-person professional services firm trains its research team on Claude Enterprise. Two weeks before, eight analysts each turn the same three anonymised client reports into a one-page brief. Time taken is recorded. A partner marks every brief against a five-point rubric covering accuracy, completeness, tone and whether anything in it is wrong. Mean time 54 minutes, mean score 3.1. Twelve weeks after training, the same eight analysts do the same task on three different reports of the same type, marked by the same partner against the same rubric. Mean time 31 minutes, mean score 3.6. Alongside it, 28-day admin data shows seven of the eight using the tool at least weekly.
What makes that defensible is boring and mechanical: the task came from the job, the rubric was written down before anyone was marked, the marker did not change, and the before exists. What would still get challenged is worth naming too. Eight people is a small sample. The second set of reports might have been easier. The partner knew which round was which, so blind the marking if you can. State those three limits yourself, in the paper, and the number gets stronger rather than weaker.
Which should you use at 20 people, and at 200?
At 20 people, use the assessed capability measure and skip the survey. At 200, run all three, because you need one number per layer and no single measure covers that many roles.
At 20, you can mark 20 pieces of work, so mark them. Usage activity across 20 people is mostly noise, since one enthusiast moves the mean several points. A survey of 20 people is 20 anecdotes with an average printed on top, and you would learn more by asking six of them properly.
At 200, marking everybody is not happening. Sample instead: 20 to 30 people, stratified across the roles that matter, chosen to include the sceptics rather than the volunteers. Usage activity starts to mean something at this size, because the variance between departments is itself the finding. The survey earns its place as a friction detector across the whole population, which is the one job it does well.
Below about 10 people, do not build a measurement framework. Ask them, watch the work, and spend the effort on the training instead.
What most teams get wrong about measuring AI adoption
The most common mistake is choosing the measure by what is easy to collect, then discovering it does not answer the question you were asked. The rest of the list is short and consistent.
Presenting adoption rate as an outcome. Adoption rate is a habit measure. It belongs in the diagnosis section, not the results section.
Setting a target before a baseline exists. "80% weekly active by Q4" is a number somebody invented in a planning meeting. Targets are only meaningful against your own starting point, on the same window.
Comparing your rate to a published benchmark. Vendor benchmarks use vendor definitions of active. You are comparing two different words.
Reading the result at 30 days and stopping. Capability and habit run on separate clocks, which is why the question of how long AI training takes has two answers rather than one.
Reporting only the average. In most teams the spread is the finding. A mean that hides two people building daily and 12 who have never logged in tells you to do nothing, which is the wrong instruction.
Not deciding in advance what result would change the plan. If no possible number would change what you do next quarter, you are collecting, not measuring, and you can stop.
All of this sits downstream of a harder question, which is what you were trying to change in the first place. If that is still open, start with how to train a team in AI and come back to measurement once there is something to measure.
Decide the measure before you book the training
The cheapest hour in an AI programme is the one spent deciding what you will measure, before anybody is trained and while a real baseline is still available. It is also the hour that almost never happens, because training feels urgent and measurement feels like something you can add later. You cannot add a baseline later.
ivee runs AI diagnostics that put a defensible number on where a team actually is before anything gets booked. Tell us what you have to report on and who is going to challenge it, and we will tell you which measure would survive that meeting. Usually it costs less than you would expect, and it is a conversation rather than a commitment.




