Measurement · Frameworks

Kirkpatrick was right. The industry stopped at level one anyway.

Four levels — reaction, learning, behaviour, results — proposed to climb from satisfaction toward impact. Seventy years on, the average program still camps at level one because climbing was always too expensive. Here’s the model, its critics, and the machinery that changes the economics.

The four levels · why programs camp at one · automating the climb

4 levelsreaction, learning, behaviour, results — the climb the model always intended
Level 1where most measurement stops: smile sheets, cheap to collect and uncorrelated with learning
Automatedlevels 2-3 become by-products when practice generates delayed measurement continuously

The model’s real problem was always the price of climbing

Level one is a survey at the door — free, instant, and famously uncorrelated with learning. Level two done honestly needs measurement at delay, which meant assessment projects nobody funded twice. Level three needs behaviour observation, which meant shadowing studies nobody funded once. So the industry standardised on the level it could afford and called the summit aspirational — a pricing failure wearing a methodology costume.

Continuous practice repriced the climb. When training runs on scheduled retrieval, level two data — knowledge at delay, per person — accumulates as a by-product; level three’s leading edge — whether knowledge holds at application depth when situations demand it — shows in scenario performance and calibration. The model didn’t need replacing; it needed an engine that made its upper floors affordable.

L1 REACTION100L2 LEARNING48L3 BEHAVIOUR32L4 RESULTS17MEASURED NOW69© 2026 FUTURE PROOF™
The industry’s Kirkpatrick pyramid as actually practised — and where automated delayed measurement moves the line. The smile-sheet research →

Level two as a standing dataset

Delayed retrieval per concept per person, error bars included — the level-two evidence that used to cost an assessment project, generated by the training itself.

100% TAUGHTL2 CONTINUOUSL2 UNMEASUREDDAY 1DAY 90© 2026 FUTURE PROOF™

Level three’s measurable leading edge

Application-depth scenarios and calibration under pressure predict transfer better than anything short of field observation — and unlike field observation, they scale.

THE BEHAVIOUR PREDICTOR THAT SCALESPERFECT© 2026 FUTURE PROOF™

Level four, honestly brokered

Results attribution stays hard; we keep it honest — measured capability as the leading indicator, your operational data closing the loop, assumptions printed throughout.

Q1Q2Q3Q4Q5L4: LEADING INDICATOR + YOUR DATA© 2026 FUTURE PROOF™

Levels three and four, instrumented

The model everyone cites, measured for once — behaviour and results made visible past the smile sheet.

Kirkpatrick — L1 to L4
L3✓
Stuck at
L1
Instrument L2 now
Proxy L3 weekly
Link L4 to ops

Interface shown as an illustration with representative numbers, not a screenshot — the layout is the product’s.

Climb past level one this quarter.

One program instrumented properly: reaction kept, learning measured at delay, behaviour’s leading edge tracked. The pyramid, finally used.

Questions buyers ask

Is the Kirkpatrick model still worth using?

As a question hierarchy, absolutely — it names the right ambitions in the right order. Its misuse (camping at level one) was economic, and the economics changed. Use the levels; automate the climb.

What do critics get right about the model?

The levels aren’t causally chained — happy learners don’t necessarily learn, and learning doesn’t guarantee transfer. Treat them as separate measurements, not a cascade, and the critique dissolves.

Should we keep collecting reaction data?

Yes, for what it measures: friction, relevance perception, delivery quality. Just stop presenting it as evidence of learning — it never was.

How does the New World Kirkpatrick version fit?

Its emphasis on leading indicators and required drivers maps neatly onto this machinery — adherence and calibration are exactly the leading indicators it asks programs to define.

What’s the minimum viable level-two upgrade?

One program’s material on scheduled retrieval with delayed measurement — a quarter later you have level-two curves and the case for instrumenting everything else.

See it on your own content.

Bring one course. We’ll show you the retention curve your current training leaves behind — and what scheduled review does to it.

  • 30 minutes, on your calendar — pick a slot here
  • Run on your own content wherever possible, not a canned deck
  • You see the dashboards, the learner surface and the evidence exports
  • No commitment — and pilot data stays yours either way