Kirkpatrick was right. The industry stopped at level one anyway.
Four levels — reaction, learning, behaviour, results — proposed to climb from satisfaction toward impact. Seventy years on, the average program still camps at level one because climbing was always too expensive. Here’s the model, its critics, and the machinery that changes the economics.
The model’s real problem was always the price of climbing
Level one is a survey at the door — free, instant, and famously uncorrelated with learning. Level two done honestly needs measurement at delay, which meant assessment projects nobody funded twice. Level three needs behaviour observation, which meant shadowing studies nobody funded once. So the industry standardised on the level it could afford and called the summit aspirational — a pricing failure wearing a methodology costume.
Continuous practice repriced the climb. When training runs on scheduled retrieval, level two data — knowledge at delay, per person — accumulates as a by-product; level three’s leading edge — whether knowledge holds at application depth when situations demand it — shows in scenario performance and calibration. The model didn’t need replacing; it needed an engine that made its upper floors affordable.
Level two as a standing dataset
Delayed retrieval per concept per person, error bars included — the level-two evidence that used to cost an assessment project, generated by the training itself.
Level three’s measurable leading edge
Application-depth scenarios and calibration under pressure predict transfer better than anything short of field observation — and unlike field observation, they scale.
Level four, honestly brokered
Results attribution stays hard; we keep it honest — measured capability as the leading indicator, your operational data closing the loop, assumptions printed throughout.
Levels three and four, instrumented
The model everyone cites, measured for once — behaviour and results made visible past the smile sheet.
Interface shown as an illustration with representative numbers, not a screenshot — the layout is the product’s.
Climb past level one this quarter.
One program instrumented properly: reaction kept, learning measured at delay, behaviour’s leading edge tracked. The pyramid, finally used.
Questions buyers ask
Is the Kirkpatrick model still worth using?
As a question hierarchy, absolutely — it names the right ambitions in the right order. Its misuse (camping at level one) was economic, and the economics changed. Use the levels; automate the climb.
What do critics get right about the model?
The levels aren’t causally chained — happy learners don’t necessarily learn, and learning doesn’t guarantee transfer. Treat them as separate measurements, not a cascade, and the critique dissolves.
Should we keep collecting reaction data?
Yes, for what it measures: friction, relevance perception, delivery quality. Just stop presenting it as evidence of learning — it never was.
How does the New World Kirkpatrick version fit?
Its emphasis on leading indicators and required drivers maps neatly onto this machinery — adherence and calibration are exactly the leading indicators it asks programs to define.
What’s the minimum viable level-two upgrade?
One program’s material on scheduled retrieval with delayed measurement — a quarter later you have level-two curves and the case for instrumenting everything else.
See it on your own content.
Bring one course. We’ll show you the retention curve your current training leaves behind — and what scheduled review does to it.
- 30 minutes, on your calendar — pick a slot here
- Run on your own content wherever possible, not a canned deck
- You see the dashboards, the learner surface and the evidence exports
- No commitment — and pilot data stays yours either way