© 2026 FUTURE PROOF™
The Uncomfortable Evidence · Transfer of Training

Why compliance training doesn’t transfer.

Almost everyone completes the annual course. Almost no one behaves differently because of it. Four decades of transfer-of-training research explain the gap between finishing a module and doing the job differently — and why the one-shot format is built to fail it; it is also the gap Future Proof™ was built to close.

TL;DR

The finding: Completing training and changing behavior are different outcomes, and the field has known it for forty years. Baldwin and Ford’s foundational review estimated that only a small fraction of what is trained is ever transferred to the job, and the meta-analyses since have confirmed that reaction and even knowledge measures predict on-the-job behavior weakly. A course-completion rate is a measure of attendance, not of transfer.

The mechanism: Transfer is not mainly a property of the course. It is moderated by trainee characteristics, training design, and — heavily — the work environment: whether there is support, opportunity to use the skill, and reinforcement over time. A single annual session gives the design levers that matter most, spacing and post-training support, almost nothing to work with.

The product: Future Proof closes the completion-to-behavior gap with spaced reinforcement scheduled after the course and retention measurement that reports what learners can still do weeks later — not whether they clicked “finish.”

In this article

  1. 01Completion is not the outcome
  2. 02How little transfers
  3. 03What the meta-analyses actually moderate
  4. 04Why the annual one-shot underdelivers
  5. 05Design features that move transfer
  6. 06Transfer is a process, not a verdict
  7. 07What the evidence doesn’t show
© 2026 FUTURE PROOF™
The route. 7 sections, from “Completion is not the outcome” to “What the evidence doesn’t show”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Corporate training is a market measured in hundreds of billions of dollars a year. The single largest slice of it is mandatory — training people take because a regulator, an insurer, or a policy requires it. That makes this article’s question one of the most economically weighty in applied psychology, and one of the least often asked out loud. After the completion certificate is issued, does anyone behave differently? The research answer has been stable for four decades, and it is not the answer the dashboard implies.

Every organization of any size runs the same ritual once a year. Employees are assigned a compliance course — anti-harassment, data security, anti-bribery, safety — they click through it, they pass a short quiz, and a dashboard turns green. Completion hits ninety-something percent. The obligation is discharged. And then, in the great majority of cases, nothing about how people actually behave at work changes.

This is not cynicism; it is the finding of one of the oldest and best-established literatures in applied psychology. The field calls the phenomenon transfer of training: how far what is learned in training carries to the job and lasts over time. The uncomfortable news has been in print since the 1980s. Transfer is the exception, not the rule. And the format most compliance training uses — a single annual session measured by completion — is close to a worst case for producing it. Future Proof™ exists in the gap between the two, but the gap itself belongs to the evidence.

Completion is not the outcome

The cleanest way to see the problem is Donald Kirkpatrick’s four levels (Kirkpatrick, 1996), the framework nearly every L&D team already uses whether or not they name it. The levels: trainees’ reactions (did they like it), learning (did they acquire the knowledge), behavior (did they apply it on the job), and results (did the organization benefit). Compliance dashboards measure something below even level one — mere completion. The mistake baked into the ritual is treating a level-one-or-below number as if it were evidence about level three.

It is not, and we can put numbers on how badly it fails. Alliger and colleagues’ meta-analysis of the Kirkpatrick criteria found that trainees’ affective reactions — how much they liked the training — were essentially uncorrelated with how much they actually learned or with later job behavior (Alliger et al., 1997). A happy sheet tells you almost nothing about transfer. Worse, even the level that seems safest — did they learn it — is a weak proxy for behavior. Passing the end-of-module quiz demonstrates that the knowledge was available in the moment, in the same context, minutes after exposure. Transfer asks a different question: can the person retrieve and apply it, weeks later, in the messy context of real work.

How little transfers

The founding statement of the problem is Baldwin and Ford’s 1988 review in Personnel Psychology, still among the most cited papers in the training literature (Baldwin & Ford, 1988). They did two things. First, they organized the whole question into a model with three classes of input — trainee traits, training design, and work environment. Those inputs feed two conditions of transfer: generalization of learning to the job, and maintenance of it over time. Second, they delivered the sobering estimate the field has repeated ever since. Only a small fraction of the investment in training produces lasting behavior change back on the job.

That framing recasts the whole enterprise, and it has held up as the field’s organizing scheme for nearly four decades. If transfer depends on trainee traits, design, and the work environment, then course content — the part compliance vendors compete on — is one lever among several. It is not obviously the decisive one. A brilliant module dropped into an unsupportive environment, reviewed once and never again, is set up to fail no matter how good the module is.

100% 67% 33% 0% End of course Week 4 Month 3 Next annual one-shot annual course reinforce spaced reinforcement transfer gap Retained behavior © 2026 FUTURE PROOF™
Figure 1. The one-shot annual course (red, dashed) decays toward baseline long before the next cycle; spaced reinforcement (teal, solid) interrupts the decline at each review. Schematic — illustrates the maintenance problem described by Baldwin & Ford (1988) and Blume et al. (2010), not a single measured dataset. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What the meta-analyses actually moderate

Twenty-two years after Baldwin and Ford framed the question, Blume, Ford, Baldwin, and Huang tested it against the piled-up data. Their 2010 meta-analysis in the Journal of Management pooled the transfer literature to ask which of those input classes actually predicts transfer, and by how much (Blume et al., 2010). Three results from that synthesis should reshape how anyone designs mandatory training.

First, the predictors of transfer are not mainly about the course. What mattered: trainee cognitive ability, personality — notably conscientiousness — motivation to transfer, and, critically, a supportive transfer climate at work. All showed meaningful links with transfer. The environment the trainee returns to is doing real work in every model. That is exactly what a maintenance-over-time view of the problem would predict.

Second, the strength of a transfer finding depends heavily on how transfer was measured. Studies that measured in the same context as the training, and right away, reported stronger transfer — systematically — than studies that measured behavior later and in a genuinely different setting (Blume et al., 2010). In plain terms: the further the measurement gets from the classroom, in time and in setting, the more the apparent effect shrinks. The end-of-module quiz is the most flattering possible test.

The catch

Where and when you measure decides what you see: same-context, immediate measures report systematically stronger transfer than delayed, different-context ones. A program that only ever runs the in-context quiz has chosen the instrument that cannot deliver bad news (Blume et al., 2010).

Third — and this is the design lever the literature keeps handing us — how training is built and delivered moderates outcomes. Arthur and colleagues pooled 1,152 effect sizes across the organizational training literature. Effectiveness varied a great deal with design and evaluation features, not just with content (Arthur et al., 2003). Training worked, on average. It worked more when the design matched the skill being trained and the yardstick being measured. Delivery format is not a neutral wrapper around content; it is part of the intervention.

The number

1,152 Effect sizes pooled across the organizational training literature: training works, on average — and works more when the design matches the skill being trained and the criterion being measured (Arthur et al., 2003).

Measured — what the dashboard records Completion recorded 90%+ 0 50% 100% of those assigned Rank order, not measured — apparent transfer by how it was measured Same context, immediate Same context, weeks later Different context, weeks latermost flattering the honest one apparent transfer effect, rank order only © 2026 FUTURE PROOF™
Figure 2. The instrument decides the answer. Top, the only number the compliance ritual reliably produces: completion, which lands in the nineties and is drawn here to scale against the whole assigned population. Bottom, what the meta-analysis found about the numbers that would actually matter — studies measuring transfer in the training context and right away report systematically stronger effects than studies measuring behavior weeks later in a genuinely different setting, which makes the end-of-module quiz the most flattering possible test (Blume et al., 2010). The two strips share no scale: the top one is a percentage of everyone assigned, the lower one a rank order with no common unit. Only the completion bar is drawn from a reported quantity. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Transfer is not just about whether learning generalizes to the job — it must also be maintained over the time that follows training.

Paraphrasing Baldwin & Ford, Personnel Psychology, 1988

Why the annual one-shot underdelivers

With the moderators on the table, the diagnosis of the compliance ritual stops being rhetorical. It becomes almost mechanical: walk the format past each lever the meta-analyses identified, and check what it does.

Walked past those levers, the standard compliance format looks almost adversarially designed. The one-shot annual course packs all exposure into a single sitting. It is graded at the moment of exposure, by a quiz in the same context as the instruction. It provides no structured reinforcement, no distributed practice, and usually no post-training support at work. Every moderator the meta-analyses flag as favorable is either absent or actively worked against.

The maintenance failure is the most predictable part. Baldwin and Ford named it explicitly: the amount transferred decreases over the time following training (Baldwin & Ford, 1988). It is a special case of a much broader regularity in memory. Without retrieval, learning decays — and it decays fastest right after acquisition. A course delivered in month one and never revisited until month twelve spends most of the year below any useful threshold of retained behavior. The dashboard, measured in month one, never sees the decline.

The mechanism that would stop the decline is not exotic. The same literature that documents the maintenance problem also points to its remedy. Spread the practice out; retrieve the content repeatedly over time; support its use in the actual work setting. Spaced retrieval and post-training reinforcement are, in transfer terms, the maintenance treatment. They are exactly what a single annual session cannot provide — and exactly what a system running after the course can.

Design features that move transfer

Two decades of this work converge on a small set of levers that reliably help, none of which require better slide decks:

  • Distribute the exposure. Splitting the same content across spaced sessions outperforms cramming it into one — the single most robust behavioral finding available, quantified across hundreds of comparisons in the distributed-practice literature (Cepeda et al., 2006), and the direct antidote to the annual sitting.
  • Measure behavior late and elsewhere. If the only measurement is an in-context immediate quiz, the number will flatter you. A delayed, transfer-appropriate check is the honest one — and, per Blume and colleagues, the one that predicts the job (Blume et al., 2010).
  • Build the environment, not just the course. A supportive transfer climate — opportunity to use the skill, cues and reinforcement at work — is among the strongest correlates of transfer in the meta-analytic record (Blume et al., 2010).
  • Match design to criterion. Arthur and colleagues show effectiveness rises when the training method fits the skill and the outcome being trained toward, rather than defaulting to one format for everything (Arthur et al., 2003).

Transfer is a process, not a verdict

The most useful recent development in this literature is a change of tense. The classic model treats transfer as an outcome to be predicted — inputs on the left, a transfer score on the right. Blume, Ford, Surface, and Olenick’s dynamic model redraws it as a loop that runs after the course ends (Blume et al., 2019). The trainee makes an initial attempt to use the skill at work, that attempt succeeds or fails, the result feeds back into their confidence and motivation, and the next attempt is more or less likely accordingly. Transfer, in this view, is not decided in the classroom. It is decided across the first weeks of attempts, each one nudging the trajectory up or down.

That reframing has a sharp practical edge, because loops have leverage points that verdicts do not. If early attempts are where paths diverge, then the highest-value window for support opens right after training — exactly when the traditional model has already moved on. Grossman and Salas’s review of the best-evidenced moderators lands on the same ground from the practitioner side. Transfer climate, supervisor support, opportunity to perform, and follow-up all sit in the post-course period (Grossman & Salas, 2011). The course sets the ceiling; the weeks after it set the number. A company that spends its whole budget on the sitting and nothing on the aftermath has funded the half of the process with less leverage.

Why it matters

If transfer is a loop of early attempts rather than a classroom verdict, the highest-value support window opens immediately after training — exactly when the traditional model has moved on. Climate, supervisor support, opportunity to perform, and follow-up all live in that post-course period (Blume et al., 2019) (Grossman & Salas, 2011).

What the evidence doesn’t show

This literature is strong, and it is easy to overclaim from. Four honest limits keep it in bounds:

  • Training is not worthless. The headline is not that training fails. Arthur and colleagues found that organizational training produces real, moderate-to-large effects on average (Arthur et al., 2003). The problem is the gap between that potential and the one-shot compliance format, not training as such.
  • The famous small number is an estimate, not a constant. Baldwin and Ford’s figure for how little transfers is a characterization of a varied literature, not a measured universal rate (Baldwin & Ford, 1988). Later reviews have argued the picture is more nuanced and, in places, more optimistic (Ford et al., 2018). Treat it as direction, not decimal.
  • Correlation, mostly. Much of the moderator evidence — transfer climate, motivation, conscientiousness — is correlational. Blume and colleagues are careful that these are associations; the causal weight of any single lever, and how they interact, is still debated (Blume et al., 2010).
  • Compliance-specific evidence is thinner. Most transfer research studies skills and knowledge broadly; rigorous behavioral evaluations of compliance courses specifically are comparatively rare. Sexual-harassment training reviews, for instance, report mixed and sometimes null effects on the behaviors that matter most (Roehling & Huang, 2018). The general model is well supported; the compliance-specific effect sizes are less settled.

Where the evidence stops

  1. 1Training is not worthless
  2. 2The famous small number is an estimate, not a constant
  3. 3Correlation, mostly
  4. 4Compliance-specific evidence is thinner
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

None of that rescues the annual one-shot. It sharpens the claim: the failure is not inevitable, and it is not the fault of training in general. It is the fault of a format that ignores every design and maintenance lever the evidence has spent forty years identifying.

There is also a governance point hiding in this literature that deserves to be said plainly. The annual one-shot persists not because anyone believes it changes behavior, but because it discharges an obligation. The completion record is the deliverable, and the record is real even when the learning is not. That position is defensible only as long as nobody claims otherwise.

The moment a company asserts that its training program manages a risk — that employees can spot the phishing attempt, apply the safety procedure, escalate the conflict — it has made an empirical claim about level-three behavior. The completion dashboard is not evidence for it (Alliger et al., 1997). The transfer literature’s real gift to compliance teams is a vocabulary for the difference. Attendance is documented; behavior is measured; and only one of those was ever the point.

Applied at Future Proof

How Future Proof™ applies this — closing the completion-to-behavior gap.

The transfer literature names two interventions the one-shot format cannot deliver, and Future Proof runs both automatically. Spaced reinforcement schedules short retrieval touches after the course — distributed across the weeks and months when the annual module would otherwise be decaying toward baseline — so maintenance stops being left to chance. Retention measurement replaces the completion tick with a delayed, transfer-appropriate check: the report says what a learner can still do weeks later and in a different context, which is the measure Blume and colleagues found actually tracks the job — not whether they clicked “finish.” Completion was never the outcome. This is the outcome, measured.

See spaced reinforcement
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Research Library PDF.

The evidence, by year

  • 1988Baldwin
  • 1996Kirkpatrick
  • 1997Alliger
  • 2003Arthur
  • 2006Cepeda
  • 2010Blume
  • 2011Grossman
  • 2018Ford
  • 2018Roehling
  • 2019Blume
© 2026 FUTURE PROOF™
The evidence base. The 10 sources cited here span 1988–2019, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Alliger, G.M., Tannenbaum, S.I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology 50(2): 341–358. DOIPDF
  2. Baldwin, T.T., & Ford, J.K. (1988). Transfer of Training: A Review and Directions for Future Research. Personnel Psychology 41(1): 63–105. DOIPDF
  3. Blume, B.D., Ford, J.K., Baldwin, T.T., & Huang, J.L. (2010). Transfer of Training: A Meta-Analytic Review. Journal of Management 36(4): 1065–1105. DOIPDF
  4. Arthur, W., Bennett, W., Edens, P.S., & Bell, S.T. (2003). Effectiveness of Training in Organizations: A Meta-Analysis of Design and Evaluation Features. Journal of Applied Psychology 88(2): 234–245. DOIPDF
  5. Ford, J.K., Baldwin, T.T., & Prasad, J. (2018). Transfer of Training: The Known and the Unknown. Annual Review of Organizational Psychology and Organizational Behavior 5: 201–225. DOIPDF
  6. Roehling, M.V., & Huang, J. (2018). Sexual harassment training effectiveness: An interdisciplinary review and call for research. Journal of Organizational Behavior 39(2): 134–150. DOIPDF
  7. Grossman, R., & Salas, E. (2011). The transfer of training: What really matters. International Journal of Training and Development 15(2): 103–120. DOIPDF
  8. Kirkpatrick, D.L. (1996). Great ideas revisited: Techniques for evaluating training programs. Training & Development 50(1): 54–59. PDF
  9. Blume, B.D., Ford, J.K., Surface, E.A., & Olenick, J. (2019). A dynamic model of training transfer. Human Resource Management Review 29(2): 270–283. DOIPDF
  10. Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132(3): 354–380. DOIPDF
Close the transfer gap

Your completion rate is green. Your transfer rate is unknown.

Book a 20-minute demo using your team’s actual compliance content. We’ll schedule the spaced reinforcement after the course and show you the retention report — what learners can still do weeks later, in a different context — instead of the completion tick that never told you anything.

10 citations Reviewed August 2026 Open peer review welcomed