Measurement · Method

How to measure training effectiveness without fooling yourself.

Most effectiveness measurement fails in one of two ways: it measures at the wrong time (course end, when memory is at its peak) or it measures the wrong thing (satisfaction, completion, hours). Here is a method that survives contact with a sceptical CFO — and the tooling that automates it.

Measure at delay · build the counterfactual · report capability

At delaythe only honest timing: what remains weeks after the course is the training’s actual product
2 linesevery effectiveness claim needs a counterfactual — the measured curve against the no-practice projection
0weight given to satisfaction scores: enjoyment and learning are uncorrelated in the research

The method, in three commitments

First commitment: measure at delay. An end-of-course score measures short-term memory in a supportive context — it flatters everyone and predicts little. Effectiveness is retrieval performance at two weeks, six weeks, a quarter. Second commitment: measure against a counterfactual. A retention number alone is uninterpretable; the claim is the gap between what your program held and what decay would have taken. Put both lines on the chart or the number is decoration.

Third commitment: report capability, not activity. Hours, completions and satisfaction go in an appendix if anywhere — the four levels people cite from Kirkpatrick were always meant to climb toward behaviour and results, and the industry’s habit of stopping at level one is how “smile sheets” became a research punchline. What a workforce can demonstrably do, at delay, per topic: that’s the report.

COURSE-END SCORE 92%SCORE AT DELAY 38%COURSE END6 WEEKS© 2026 FUTURE PROOF™
Commitment one and three in a single chart: the number measured at course end versus the number measured at delay. Only one predicts behaviour on the job. Why completion metrics mislead →

Continuous measurement, free with practice

When training runs on scheduled retrieval, every review doubles as a delayed measurement — so the effectiveness dataset accumulates continuously, with no assessment days and no survey fatigue.

100% TAUGHTTWO LINES, ONE CLAIMNO COUNTERFACTUALDAY 1DAY 90© 2026 FUTURE PROOF™

Depth as a dimension

A program can raise recall and leave application untouched. Measuring per Bloom level catches the difference between knowing the rule and using it under pressure — two different effectiveness claims.

REACTION100RECALL77APPLICATION51BEHAVIOUR36RESULTS23© 2026 FUTURE PROOF™

Uncertainty stated, always

Small cohorts and short windows widen error bars; a serious measurement shows them. One score is not a fact — the platform prints intervals so decisions inherit honest confidence.

WEEK 2: WIDE BARSQUARTER: NARROWEARLY DATAQUESTION 24© 2026 FUTURE PROOF™

Effectiveness, past level two

Reaction scores flatter; retention curves testify — measurement that survives the question ‘but did it work?’

Effectiveness — 4 measures
L3
Surveys say
4.6★
Baseline first
Measure at 90d
Report the delta

Interface shown as an illustration with representative numbers, not a screenshot — the layout is the product’s.

Get the method as a running system.

The platform operationalises all three commitments out of the box — measurement at delay, counterfactual charts, capability reporting.

Questions buyers ask

What about measuring business impact directly?

Aim for it where the causal chain is short — error rates after procedure training, say. But impact metrics are noisy and confounded; measured capability at delay is the reliable leading indicator you can attribute. Report both, trust the attribution chain in that order.

Is Kirkpatrick wrong?

The model’s ambition — climb from reaction to results — is right; the industry’s practice of camping at level one is what fails. The commitments here are a way of taking levels two and three seriously enough to automate them.

How do we run a counterfactual without denying training to a control group?

Use projection: decay curves from your own early data provide the no-practice baseline without withholding anything. Where natural comparisons exist — staggered rollouts, late cohorts — use them opportunistically.

How soon after rollout can effectiveness be claimed?

Cohort-level curves stabilise within a few review cycles — weeks, not quarters. Claim early results with their error bars and let the intervals narrow as data accumulates.

Can this method evaluate our existing non-adaptive training?

Yes — run the measurement layer on any program’s output. That baseline, incidentally, is the strongest argument you’ll ever have for changing the program.

See it on your own content.

Bring one course. We’ll show you the retention curve your current training leaves behind — and what scheduled review does to it.

  • 30 minutes, on your calendar — pick a slot here
  • Run on your own content wherever possible, not a canned deck
  • You see the dashboards, the learner surface and the evidence exports
  • No commitment — and pilot data stays yours either way