Back

AI Skills Assessment: 8 Things to Map Before Training

A baseline you can take in a week, so the training gets designed for the team you have rather than the average of it.

Date

Reading time

11

min

Amelia Miller

Co-founder and CEO

An AI skills assessment is a baseline reading of what a team can already do with AI tools, taken before any training is booked, so the programme can be designed for the people in the room rather than for an imagined average.

Assess eight things: which tools each person can actually reach, what they did with AI in the last 30 days, prompting capability tested on one real task, which work is still manual, appetite kept separate from ability, what people believe the data rules are, who the internal champion would be, and how far self-ratings sit from tested results.

Most assessments come back as a single number for the whole team. That number is the least useful thing you can collect, because the decision it is supposed to inform, which is whether you run one training track or two, depends entirely on the shape of the distribution it just flattened.

What makes something worth assessing before AI training?

Something is worth assessing if the answer changes a decision you are about to make. That is the whole filter, and it removes most of what appears on competitor assessment templates.

Three criteria, applied in order:

  • It changes a decision. Track structure, content, sequencing, format, or who goes first. If every possible answer leads to the same programme, do not collect it.

  • One person can capture it in under a week, without a budget. Anything needing a procurement cycle is a different project.

  • It can come back badly. A question nobody can fail is a marketing exercise. If there is no answer that would make you change the plan, you have written a lead magnet.

Everything below clears all three. Note what is missing: no maturity level, no five-point scale, no composite score. Those are outputs of an assessment, not the assessment, and rolling eight readings into one percentage destroys the information you spent the week collecting.

The 8 things to assess before you train a team in AI

These eight are ordered by how cheap they are to collect, not by how important they are. Items 1 and 2 come out of a system you already pay for; items 3 and 8 cost the most and change the most.

1. Which AI tools each person can actually reach today

Licence entitlement and real access are different things, and the gap between them is where a training day goes to die. Someone with a Microsoft 365 Copilot licence who cannot reach the SharePoint site their work lives in has a tool that demos well and does nothing.

Capture it from the admin centre rather than by asking. Microsoft 365 Copilot reports enabled and active users per person, and Claude Enterprise, ChatGPT Enterprise and Google Workspace expose equivalents. It takes an afternoon.

The decision it changes: whether you are training people on a tool they hold, or running an access-remediation project with a training day stapled to the end.

2. What people did with AI in the last 30 days

Behaviour tells you where the team starts; opinion tells you how they feel about it. Pull the previous 28 or 30 days of activity from the admin centre before anyone knows an assessment is happening, because announcing it moves the number.

You are looking for two things: how many people touched the tool at all, and how concentrated the usage is. Ten people using it daily and forty who opened it once is a completely different starting point from fifty people using it weekly, even where the aggregate looks similar.

The decision it changes: whether the programme's job is to create first use or to deepen existing use. Those need different content and different lengths. What you do with this number after the training is a separate question from what it tells you now.

3. Prompting capability, tested on one real piece of work

Self-declared prompting skill is close to worthless, and this is the item people most often skip because testing feels heavy-handed. It does not have to be. Give eight to ten people one task from their actual job, thirty minutes, and the tool they already have.

Read the outputs yourself. You are not marking; you are sorting into three piles: could not get started, got something usable, got something better than they would have produced alone. Those three piles are your tracks.

The decision it changes: the level you pitch at, and whether the middle pile is large enough to justify a single cohort.

4. Which work is still manual, repetitive and already reviewed by a human

Training with no target work attached does not transfer, and the safest first targets share three properties: they happen often, they are low-stakes, and somebody already checks them. Ask each line manager for the three tasks their team does most often that nobody enjoys.

Repetition gives people enough reps to build a habit. Existing human review means a mistake gets caught for free, which is what makes it a reasonable place to let people be bad at something for a fortnight.

The decision it changes: what the training is actually about. A programme built on the finance team's month-end reconciliation is a different product from one built on generic prompt craft.

5. Appetite, asked separately from ability

Appetite and ability are independent, and conflating them into one "readiness" question produces a number that means nothing. A capable sceptic and an enthusiastic beginner need opposite things: one needs evidence, the other needs guard rails.

Two questions, anonymous, on a scale: how useful do you expect this to be for your work, and how comfortable are you being seen getting it wrong. The second one predicts more than the first.

The decision it changes: who goes in the first cohort. Put the capable sceptics in early if they will engage, because they are the ones colleagues believe afterwards.

6. What people believe they are allowed to do with company data

Ask what people think the rules are, not what the rules say. The distance between the two is the real finding, and it is usually larger than anyone in the room expects.

One anonymous question does it: what would you do if you needed to summarise a client document tomorrow. If the answers include a personal ChatGPT account, you have a governance finding rather than a training finding, and it needs settling before the training rather than after.

The decision it changes: whether the first session is about capability or about what happens to inputs. Order matters, because teaching people to be faster at something they should not be doing is an expensive mistake.

7. Who would be the internal champion, by name

Every programme that survives has one person inside the team who answers questions in the fortnight after the training ends. If you cannot name them before you start, the programme has no owner, only a sponsor.

Find them from item 2 rather than by asking for volunteers. The person already using the tool most, who is not the most senior person in the room, is usually the right answer.

The decision it changes: whether you are buying training or buying training plus twelve weeks of follow-up. It also changes who needs to be in the design conversation.

8. How far self-ratings sit from tested results

This is the one nobody collects, and it is close to free once you have items 3 and 5. Ask people to rate themselves before the task in item 3, then compare the rating against what they actually produced.

The direction of the error matters more than its size. Confident people who cannot do the task will teach colleagues the wrong thing, at speed. Capable people who rate themselves low will sit out a programme they would have led.

The decision it changes: how much you trust every other self-reported number in the assessment, including the ones you have already collected.

Why the spread inside a team matters more than the average

The average tells you nothing about whether one training track will work, and the spread tells you almost everything. A team containing solution architects building on Claude daily and colleagues who have spent a career on legacy systems has an average that describes nobody in it.

Self-rating cannot give you that spread, and the reason is not the one usually offered. Workera compared self-assessments against adaptive tests across more than 22,000 domain assessments taken through early 2024 by several thousand of its users, and found 56% underestimated their own level, 32% overestimated, and 11% landed within 10 points of their tested result on a 300-point scale. Two caveats worth stating: those were general skill domains rather than AI specifically, and Workera sells assessments.

The received wisdom is that people talk themselves up. In that dataset the larger error runs the other way. Either way the practical consequence is the same, and it is worse than inflation: self-rating scrambles the distribution rather than shifting it, so the people you most need to identify at both ends are the ones the survey is least likely to place correctly.

There is a regulatory version of this argument too. Article 4 of the EU AI Act, which has applied since 2 February 2025, requires deployers to ensure a sufficient level of AI literacy among staff "taking into account their technical knowledge, experience, education and training and the context the AI systems are to be used in". That is a legal instruction to account for variation. A single team-wide score cannot demonstrate you did.

Role variation is measurable at national scale, which is a reasonable prior for what sits inside a mixed team. The UK government's assessment of AI capabilities and the UK labour market, drawing on IMF estimates, splits the UK workforce into roughly equal thirds: 35% in occupations with high AI exposure and high complementarity, 32% high exposure and low complementarity, and 33% low exposure. Three populations, one workforce.

How do you test AI skills instead of asking about them?

Give people a real task, watch what comes back, and sort the results. Testing at baseline is a sorting exercise, not a marking exercise, and treating it as the latter is why most teams decide it is too much work and send a survey instead.

Use work the team already does, not a puzzle. Three anonymised client documents to summarise, ten historic support tickets to draft replies to, one messy spreadsheet to reconcile. Thirty minutes, their own tool, no preparation.

A scored rubric with a named marker is the right instrument later, when you need a number that survives a board question rather than a decision about cohorts. That is a different job with a different cost, and the case for an assessed capability measure sets out what it takes. For individuals who want a sense of where they personally stand before any of this, ivee's AI fluency quiz does that job and does not need rebuilding at team level.

How long does an AI skills assessment take?

Five to eight working days of one person's part-time attention, for a team of 30 to 60. Most of that is waiting for survey responses and line-manager conversations rather than analysis.

The rough shape: half a day pulling admin-centre data for items 1 and 2, two days for the anonymous survey covering items 5, 6 and 8 to come back, half a day setting the task in item 3, two days for people to do it, and a day reading outputs and writing up. Line-manager conversations for item 4 run alongside.

Two things stretch it. Admin access you do not have adds a week, sometimes two, because it is an IT request rather than a task. And a team above roughly 100 people needs sampling rather than a census, which adds a design decision at the front but does not add much elapsed time.

What it should produce is one page: three piles of people, the three target tasks, the named champion, and the one thing that would block the programme. If the output is longer than that, it is a report rather than a decision. Training your team in AI is where that page turns into a programme.

What happens if you assess and then do nothing

Assessing a team and then not acting damages trust more than never asking would have. People say what they cannot do, watch nothing change, and correctly conclude that the next survey is not worth their time.

The failure is common because assessment is cheap and acting on it is not. A finding that half the team cannot reach the tool implies an IT project. A finding that the confident people are wrong implies telling them. Both are harder than the assessment that produced them, so both get deferred, and the assessment becomes the deliverable by default.

Two protections, decided before the first question goes out. Settle which finding would make you tear up the plan, and who holds the authority to do it, so that the answer is not discovered later to be nobody. And tell people what you will do with the findings when you ask, then do that within a month, even where the answer is that nothing changes and here is why.

Where an assessment shows the team is not ready to be trained at all, the correct response is to say so and fix the blocker first. That is a better outcome than a training day booked to clear a budget line.

Find out where your team actually starts

ivee's AI diagnostic exists for the team that already suspects it is uneven and cannot yet describe how. The finding clients least expect is item 8, because it tells them which of the numbers they already report to stop trusting. Send us the programme you were about to book and who it is for, and we will tell you what to measure first, and whether you need us at all.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.