Research · The Uncomfortable Evidence
The Uncomfortable Evidence · Transfer of Training

Why compliance training doesn’t transfer.

Almost everyone completes the annual course. Almost no one behaves differently because of it. Four decades of transfer-of-training research explain the gap between finishing a module and doing the job differently — and why the one-shot format is built to fail it; it is also the gap Future Proof™ was built to close.

TL;DR

The finding: Completing training and changing behavior are different outcomes, and the field has known it for forty years. Baldwin and Ford’s foundational review estimated that only a small fraction of what is trained is ever transferred to the job, and the meta-analyses since have confirmed that reaction and even knowledge measures predict on-the-job behavior weakly. A course-completion rate is a measure of attendance, not of transfer.

The mechanism: Transfer is not mainly a property of the course. It is moderated by trainee characteristics, training design, and — heavily — the work environment: whether there is support, opportunity to use the skill, and reinforcement over time. A single annual session gives the design levers that matter most, spacing and post-training support, almost nothing to work with.

The product: Future Proof closes the completion-to-behavior gap with spaced reinforcement scheduled after the course and retention measurement that reports what learners can still do weeks later — not whether they clicked “finish.”

Every organization of any size runs the same ritual once a year. Employees are assigned a compliance course — anti-harassment, data security, anti-bribery, safety — they click through it, they pass a short quiz, and a dashboard turns green. Completion hits ninety-something percent. The obligation is discharged. And then, in the great majority of cases, nothing about how people actually behave at work changes.

This is not cynicism; it is the finding of one of the oldest and best-established literatures in applied psychology. The field calls the phenomenon transfer of training: the degree to which what is learned in training is generalized to the job and maintained over time. The uncomfortable news, which has been in print since the 1980s, is that transfer is the exception rather than the rule, and that the format most compliance training uses — a single annual session measured by completion — is close to a worst case for producing it. Future Proof™ exists in the gap between the two, but the gap itself belongs to the evidence.

Completion is not the outcome

The cleanest way to see the problem is Donald Kirkpatrick’s four levels, the framework nearly every L&D team already uses whether or not they name it: trainees’ reactions (did they like it), learning (did they acquire the knowledge), behavior (did they apply it on the job), and results (did the organization benefit). Compliance dashboards measure something below even level one — mere completion. The mistake baked into the ritual is treating a level-one-or-below number as if it were evidence about level three.

It is not, and we can put numbers on how badly it fails. Alliger and colleagues’ meta-analysis of the Kirkpatrick criteria found that trainees’ affective reactions — how much they liked the training — were essentially uncorrelated with how much they actually learned or with later job behavior (Alliger et al., 1997). A happy sheet tells you almost nothing about transfer. Worse, even the level that seems safest — did they learn it — is a weak proxy for behavior. Passing the end-of-module quiz demonstrates that the knowledge was available in the moment, in the same context, minutes after exposure. Transfer asks a different question: can the person retrieve and apply it, weeks later, in the messy context of real work.

How little transfers

The foundational statement of the problem is Baldwin and Ford’s 1988 review in Personnel Psychology, still among the most cited papers in the training literature (Baldwin & Ford, 1988). They did two things. First, they organized the entire question into a model with three classes of input — trainee characteristics, training design, and work environment — feeding two conditions of transfer: generalization of learning to the job and maintenance of it over time. Second, they delivered the sobering estimate the field has repeated ever since: that only a small fraction of the investment in training produces lasting behavioral change back on the job.

That framing reframes the whole enterprise. If transfer depends on trainee characteristics, design, and the work environment, then the course content — the part compliance vendors compete on — is one lever among several, and not obviously the decisive one. A brilliant module dropped into an unsupportive environment, reviewed once and never again, is set up to fail regardless of how good the module is.

100% 67% 33% 0% End of course Week 4 Month 3 Next annual one-shot annual course reinforce spaced reinforcement transfer gap Retained behavior
Figure 1. The one-shot annual course (red, dashed) decays toward baseline long before the next cycle; spaced reinforcement (purple, solid) interrupts the decline at each review. Schematic — illustrates the maintenance problem described by Baldwin & Ford (1988) and Blume et al. (2010), not a single measured dataset.

What the meta-analyses actually moderate

Twenty-two years after Baldwin and Ford framed the question, Blume, Ford, Baldwin, and Huang tested it against accumulated data. Their 2010 meta-analysis in the Journal of Management pooled the transfer literature to ask which of those input classes actually predicts transfer, and how much (Blume et al., 2010). Three results from that synthesis should reshape how anyone designs mandatory training.

First, the predictors of transfer are not primarily about the course. Trainee cognitive ability and personality (notably conscientiousness), motivation to transfer, and — critically — a supportive transfer climate at work all showed meaningful relationships with transfer. The environment the trainee returns to is doing real work in every model, which is exactly what you would predict from a maintenance-over-time view of the problem.

Second, the strength of a transfer finding depends heavily on how transfer was measured. Blume and colleagues found that studies relying on the same measurement context as the training, and on immediate rather than delayed measures, reported systematically stronger transfer than studies that measured behavior later and in a genuinely different context (Blume et al., 2010). In plain terms: the further the measurement gets from the classroom — in time and in setting — the more the apparent effect shrinks. The end-of-module quiz is the most flattering possible test.

Third — and this is the design lever the literature keeps handing us — how training is built and delivered moderates outcomes. Arthur and colleagues’ meta-analysis of 1,152 effect sizes across the organizational training literature found that effectiveness varied substantially with training-design and evaluation features, not just with content (Arthur et al., 2003). Training worked, on average, and it worked more when the design matched the skill being trained and the criterion being measured. Delivery format is not a neutral wrapper around content; it is part of the intervention.

Transfer is not just about whether learning generalizes to the job — it must also be maintained over the time that follows training.

Paraphrasing Baldwin & Ford, Personnel Psychology, 1988

Why the annual one-shot underdelivers

Put the pieces together and the standard compliance format looks almost adversarially designed. The one-shot annual course concentrates all exposure into a single sitting; it is typically evaluated at the moment of exposure by a quiz in the same context as the instruction; and it provides no structured reinforcement, no distributed practice, and usually no post-training support in the work environment. Every moderator that the meta-analyses identify as favorable is either absent or actively worked against.

The maintenance failure is the most predictable part. Baldwin and Ford named it explicitly — the amount transferred decreases over the time following training (Baldwin & Ford, 1988) — and it is a special case of a much broader regularity in memory: without retrieval, learning decays, and it decays fastest right after acquisition. A course delivered in month one and never revisited until month twelve spends most of the year sitting below any useful threshold of retained behavior. The dashboard, measured in month one, never sees the decline.

The mechanism that would arrest the decline is not exotic. The same literature that documents the maintenance problem also points to its remedy: distribute the practice, retrieve the content repeatedly over time, and support its use in the actual work setting. Spaced retrieval and post-training reinforcement are, in transfer terms, the maintenance intervention. They are precisely what a single annual session cannot provide, and precisely what a system running after the course can.

Design features that move transfer

Two decades of this work converge on a small set of levers that reliably help, none of which require better slide decks:

  • Distribute the exposure. Splitting the same content across spaced sessions outperforms cramming it into one — the single most robust behavioral finding available, and the direct antidote to the annual sitting.
  • Measure behavior late and elsewhere. If the only measurement is an in-context immediate quiz, the number will flatter you. A delayed, transfer-appropriate check is the honest one — and, per Blume and colleagues, the one that predicts the job (Blume et al., 2010).
  • Build the environment, not just the course. A supportive transfer climate — opportunity to use the skill, cues and reinforcement at work — is among the strongest correlates of transfer in the meta-analytic record (Blume et al., 2010).
  • Match design to criterion. Arthur and colleagues show effectiveness rises when the training method fits the skill and the outcome being trained toward, rather than defaulting to one format for everything (Arthur et al., 2003).

What the evidence doesn’t show

This literature is strong, and it is easy to overclaim from. Four honest limits keep it in bounds:

  • Training is not worthless. The headline is not that training fails. Arthur and colleagues found that organizational training produces real, moderate-to-large effects on average (Arthur et al., 2003). The problem is the gap between that potential and the one-shot compliance format, not training as such.
  • The famous small number is an estimate, not a constant. Baldwin and Ford’s figure for how little transfers is a characterization of a varied literature, not a measured universal rate (Baldwin & Ford, 1988). Later reviews have argued the picture is more nuanced and, in places, more optimistic (Ford et al., 2018). Treat it as direction, not decimal.
  • Correlation, mostly. Much of the moderator evidence — transfer climate, motivation, conscientiousness — is correlational. Blume and colleagues are careful that these are associations; the causal weight of any single lever, and how they interact, is still debated (Blume et al., 2010).
  • Compliance-specific evidence is thinner. Most transfer research studies skills and knowledge broadly; rigorous behavioral evaluations of compliance courses specifically are comparatively rare. Sexual-harassment training reviews, for instance, report mixed and sometimes null effects on the behaviors that matter most (Roehling & Huang, 2018). The general model is well supported; the compliance-specific effect sizes are less settled.

None of that rescues the annual one-shot. It sharpens the claim: the failure is not inevitable, and it is not the fault of training in general. It is the fault of a format that ignores every design and maintenance lever the evidence has spent forty years identifying.

Applied at Future Proof

How Future Proof™ applies this — closing the completion-to-behavior gap.

The transfer literature names two interventions the one-shot format cannot deliver, and Future Proof runs both automatically. Spaced reinforcement schedules short retrieval touches after the course — distributed across the weeks and months when the annual module would otherwise be decaying toward baseline — so maintenance stops being left to chance. Retention measurement replaces the completion tick with a delayed, transfer-appropriate check: the report says what a learner can still do weeks later and in a different context, which is the measure Blume and colleagues found actually tracks the job — not whether they clicked “finish.” Completion was never the outcome. This is the outcome, measured.

See spaced reinforcement
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Research Library PDF.

  1. Alliger, G.M., Tannenbaum, S.I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology 50(2): 341–358. DOIPDF
  2. Baldwin, T.T., & Ford, J.K. (1988). Transfer of Training: A Review and Directions for Future Research. Personnel Psychology 41(1): 63–105. DOIPDF
  3. Blume, B.D., Ford, J.K., Baldwin, T.T., & Huang, J.L. (2010). Transfer of Training: A Meta-Analytic Review. Journal of Management 36(4): 1065–1105. DOIPDF
  4. Arthur, W., Bennett, W., Edens, P.S., & Bell, S.T. (2003). Effectiveness of Training in Organizations: A Meta-Analysis of Design and Evaluation Features. Journal of Applied Psychology 88(2): 234–245. DOIPDF
  5. Ford, J.K., Baldwin, T.T., & Prasad, J. (2018). Transfer of Training: The Known and the Unknown. Annual Review of Organizational Psychology and Organizational Behavior 5: 201–225. DOIPDF
  6. Roehling, M.V., & Huang, J. (2018). Sexual harassment training effectiveness: An interdisciplinary review and call for research. Journal of Organizational Behavior 39(2): 134–150. DOIPDF
  7. Grossman, R., & Salas, E. (2011). The transfer of training: What really matters. International Journal of Training and Development 15(2): 103–120. DOIPDF
  8. Kirkpatrick, D.L. (1996). Great ideas revisited: Techniques for evaluating training programs. Training & Development 50(1): 54–59. PDF
  9. Blume, B.D., Ford, J.K., Surface, E.A., & Olenick, J. (2019). A dynamic model of training transfer. Human Resource Management Review 29(2): 270–283. DOIPDF
  10. Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132(3): 354–380. DOIPDF
Close the transfer gap

Your completion rate is green. Your transfer rate is unknown.

Book a 20-minute demo using your team’s actual compliance content. We’ll schedule the spaced reinforcement after the course and show you the retention report — what learners can still do weeks later, in a different context — instead of the completion tick that never told you anything.

10 citations Reviewed July 2026 Open peer review welcomed