A method for proving training changed what a person can do, with Ofsted wording, provider thresholds and end-point assessment grades read for what they measure.

How to Measure Workplace Training Outcomes

Guides

By James Cotton · Last updated · 7 min read

Part of our topic guide on AI Skills for Business.

By James Cotton, Founder, iO-Sphere

Write the change down before training starts

Before the programme begins, write one sentence about a named person and a task you could watch happen. That is the moment you decide what the money is for. It is also the easiest time to define an outcome, because nobody has been trained yet and nothing needs explaining away.

Nameable changes in a data or AI role are concrete. Each of these could be written down on day one:

  • An analyst rebuilds a two-day manual report so it runs in an afternoon and anyone on the team can rerun it.
  • An operations lead reads a supplier's data-processing terms and flags the risk before signing, escalating only the unusual cases, which is the ground a data protection apprenticeship is built to cover.
  • A team member notices a model is scoring the wrong population and stops its output before it reaches a customer.

Each passes one check. Could a colleague who was not in the room tell whether it arrived? A target like "understands governance better" fails, because nobody can see it happen.

A satisfaction score cannot carry a change like this. It records how the day felt. The CIPD's Learning at Work 2023 survey (fieldwork January to February 2023, published 27 June 2023) found that 50% of respondents agreed they had a process for assessing the impact of learning.

Transfer, the move from course to job, was the weaker link. Only 7% strongly agreed they had a process for supporting learning transfer, with a further 39% agreeing. That 2023 report remains the most recent CIPD wave with published findings.

Look for it in behaviour, at Kirkpatrick Level 3

Kirkpatrick Partners set out a four-level model, created in the 1950s. Level 1 is Reaction, Level 2 is Learning, Level 3 is Behaviour and Level 4 is Results. The sentence you wrote lives at Level 3.

Their definition of Behaviour is "the degree to which the target audience performs the critical behaviors in its environment and is supported in and accountable for its performance". Note the second half. A manager who never gives the person a chance to use the skill has broken Level 3 as surely as a weak course.

Level 4 is where the business feels it. It matters, and a dozen things besides training move it. Level 3 is the working level because you can observe it directly, one short step from the results you care about. The benefits employers report in the Department for Education's employer survey are gathered in our guide to data and AI apprenticeship ROI.

The ROI Institute's Phillips ROI Methodology adds a fifth level that compares the monetary benefits of a programme with its costs. Phillips calls Level 3 "application and implementation", a useful reminder that knowing and doing are separate events.

For an employer training one or two people, Level 3 is the sensible place to stop. A fifth-level figure means isolating the training's effect from everything else and pricing it. That effort pays on a large programme across many staff. For a single analyst, an observed change in the work answers the question far more cheaply.

Checking a short course

A short course has no independent assessment at the end, so the check is yours to build. Write the change before the course runs. Then look for it in the work three or four weeks later, once the trainer has gone and the new habit has either held or faded.

Keep the target small enough to show inside that window: one report rebuilt, one prompt workflow the team now uses, one check someone runs before sending figures out. Our data and AI fluency training is designed to be judged against a target of that size, at Level 3 like everything else on this page.

When the training is an apprenticeship: inspection, thresholds and the end-point assessment

Apprenticeships come with public yardsticks a short course lacks. They do different jobs. Ofsted inspects and grades a provider's quality, while the Apprenticeship Accountability Framework flags a provider's performance to the funding body. Neither measures your person, though each is worth knowing when you choose a provider.

What Ofsted grades

In version 2.0 of its further education and skills inspection toolkit, published in June 2026, Ofsted sets out the renewed framework in use since 10 November 2025. There is no overall grade. Each evaluation area is graded on a five-point scale: urgent improvement, needs attention, expected standard, strong standard or exceptional.

The Achievement area asks whether learners and apprentices "develop knowledge, skills and professional behaviours as they progress". At the expected standard they "can apply these effectively"; at the strong standard, "expertly, fluently and automatically". Inspectors count "contribution to workplace activities" as a source of evidence.

One line matters for anyone reading a report: "Inspectors do not review providers' internal data." They build evidence from what they see. A report therefore tells you how well a provider's learners in general apply what they learn, and nothing about one person.

What the accountability framework flags

Under the Apprenticeship Accountability Framework, which gov.uk last updated on 17 February 2026 and the Department for Work and Pensions has owned since 1 April 2026, a provider is flagged on these thresholds:

  • Achievement: below 50% is "at risk"; 50 to under 60% "needs improvement".
  • Retention and withdrawals: retention below 52% is at risk and 52 to under 62% needs improvement; withdrawals of 20% or more are at risk.
  • Feedback: apprentice and employer feedback averaging below 2.5 needs improvement.

Those thresholds record who finishes and who stays. Hours belong in the same column. Advanced Data & AI, Data & AI Governance and Data & AI Strategy each require at least 370 hours of off-the-job training, according to Skills England. Hours completed prove time, not capability.

What the end-point assessment grades

The Advanced Data & AI assessment plan of 2 June 2021 hands the end-point assessment to an independent assessment organisation, so someone outside the provider tests the person. Each assessment method is graded fail, pass or distinction. The methods are weighted equally and combine into an overall fail, pass, merit or distinction.

The standard behind Data & AI Governance and Data & AI Strategy is ST0967, and its version 2.0, dated 9 September 2026, grades pass or distinction and includes at least one project. A project comes closest to the sentence you wrote at the start. It is work the person produced, judged by someone with no stake in the result.

A grade tells you what you want only if the training built toward it on real problems. On our apprenticeship routes, apprentices are assessed on real work, and the moment to watch is when set exercises stop. Early on, people practise on realistic case studies and simulations, which we call Prism, built to look like real business data and the decisions that come with it. After that they move to end-to-end real work, corrected as they do it, and the change you named starts to show in things the business uses.

The review at twelve months

A year in, sit down with the person and their line manager, with the sentence from the start in hand. Four questions do most of the work:

  1. What is the hardest real task this person now handles without a safety net?
  2. What have they produced that the team relies on?
  3. Which call did they make this year that they would have escalated, or missed, a year ago?
  4. If they moved team tomorrow, what would their replacement have to be taught first?

If the answers point at the sentence, you have your outcome at Level 3. If they point somewhere else, write that down as well, since useful capability often arrives next to the one you planned. For senior staff on strategy routes, the questions lean toward judgement, as our guide to AI strategy programme outcomes explains.

An analyst's answers change stage by stage, and the data analyst apprentice employer guide sets out what to expect at each stage. Of the four, one question decides the review. What is the hardest real task this person now handles without a safety net?

Frequently asked questions

Can we set the target after the programme has started?

You can, and it beats setting none. Write it with the person and their manager, about work they are doing now, and date it. The risk is fitting the target to whatever has already changed, so choose something that has not happened yet and that someone outside the team would notice.

How do we measure a team rather than one person?

Write one sentence for each person, then add one for the team: a process the whole group now runs differently, or a piece of work that used to depend on a single expert and no longer does. The individual sentences show who changed. The team sentence shows whether the change spread, which is usually what a manager sending several people wanted.

Who should write the target sentence?

The line manager and the person together, before the first session. The manager knows which task matters to the business; the person knows where they get stuck. A target written by one without the other tends to miss, either aiming at work the person never touches or at a skill nobody will ask them to use.

Will it fit your cohort?

Tell us your cohort size and what your people need to be doing by next year, and we will say whether we are the right shape, including when we are not.