How to measure training effectiveness without fooling yourself.
Most effectiveness measurement fails in one of two ways: it measures at the wrong time (course end, when memory is at its peak) or it measures the wrong thing (satisfaction, completion, hours). Here is a method that survives contact with a sceptical CFO — and the tooling that automates it.
The method, in three commitments
First commitment: measure at delay. An end-of-course score measures short-term memory in a supportive context — it flatters everyone and predicts little. Effectiveness is retrieval performance at two weeks, six weeks, a quarter. Second commitment: measure against a counterfactual. A retention number alone is uninterpretable; the claim is the gap between what your program held and what decay would have taken. Put both lines on the chart or the number is decoration.
Third commitment: report capability, not activity. Hours, completions and satisfaction go in an appendix if anywhere — the four levels people cite from Kirkpatrick were always meant to climb toward behaviour and results, and the industry’s habit of stopping at level one is how “smile sheets” became a research punchline. What a workforce can demonstrably do, at delay, per topic: that’s the report.
Continuous measurement, free with practice
When training runs on scheduled retrieval, every review doubles as a delayed measurement — so the effectiveness dataset accumulates continuously, with no assessment days and no survey fatigue.
Depth as a dimension
A program can raise recall and leave application untouched. Measuring per Bloom level catches the difference between knowing the rule and using it under pressure — two different effectiveness claims.
Uncertainty stated, always
Small cohorts and short windows widen error bars; a serious measurement shows them. One score is not a fact — the platform prints intervals so decisions inherit honest confidence.
Effectiveness, past level two
Reaction scores flatter; retention curves testify — measurement that survives the question ‘but did it work?’
Interface shown as an illustration with representative numbers, not a screenshot — the layout is the product’s.
Get the method as a running system.
The platform operationalises all three commitments out of the box — measurement at delay, counterfactual charts, capability reporting.
The evidence this page stands on
Questions buyers ask
What about measuring business impact directly?
Aim for it where the causal chain is short — error rates after procedure training, say. But impact metrics are noisy and confounded; measured capability at delay is the reliable leading indicator you can attribute. Report both, trust the attribution chain in that order.
Is Kirkpatrick wrong?
The model’s ambition — climb from reaction to results — is right; the industry’s practice of camping at level one is what fails. The commitments here are a way of taking levels two and three seriously enough to automate them.
How do we run a counterfactual without denying training to a control group?
Use projection: decay curves from your own early data provide the no-practice baseline without withholding anything. Where natural comparisons exist — staggered rollouts, late cohorts — use them opportunistically.
How soon after rollout can effectiveness be claimed?
Cohort-level curves stabilise within a few review cycles — weeks, not quarters. Claim early results with their error bars and let the intervals narrow as data accumulates.
Can this method evaluate our existing non-adaptive training?
Yes — run the measurement layer on any program’s output. That baseline, incidentally, is the strongest argument you’ll ever have for changing the program.
See it on your own content.
Bring one course. We’ll show you the retention curve your current training leaves behind — and what scheduled review does to it.
- 30 minutes, on your calendar — pick a slot here
- Run on your own content wherever possible, not a canned deck
- You see the dashboards, the learner surface and the evidence exports
- No commitment — and pilot data stays yours either way