© 2026 FUTURE PROOF™
The Uncomfortable Evidence · Compliance Training

Mandatory training: what billions of hours buy.

Compliance, ethics and security training may be the largest forced-participation education programme on earth — and almost the only thing it measures about itself is completion. The outcome evidence is uncomfortable: small, design-dependent effects at best, and in the best security experiment ever run inside a company, none at all. The fix has been sitting in the learning literature for decades.

TL;DR

The finding: The measured record of mandatory training is thin relative to its scale. Meta-analyses of ethics instruction find small, strongly design-dependent effects (Waples et al., 2009; Medeiros et al., 2017). The largest in-company phishing experiment found that embedded training — the industry’s default teachable-moment product — did not reduce employees’ likelihood of falling for phishing, and by one long-run analysis increased it (Lain et al., 2022). The annual fire-hose format fails in exactly the ways the learning literature predicts.

The mechanism: The standard module violates the three best-established results in learning science: massed single sessions lose to spaced practice (Cepeda et al., 2006), knowledge transfer does not equal behaviour change (Bada et al., 2019), and whether training reaches the job depends on practice, follow-up and the work environment — none of which a seat-time module supplies (Blume et al., 2010). Completion, the metric nearly every programme reports, sits below the lowest rung of any serious evaluation framework (Kirkpatrick & Kirkpatrick, 2006).

The product: Future Proof™ replaces the annual fire-hose with what the evidence supports — spaced micro-boosters scheduled against forgetting, scenario practice with feedback, just-in-time delivery at the moment of need, and behavioural measurement in place of completion dashboards.

In this article

  1. 01The largest training programme on earth
  2. 02What the ethics meta-analyses found
  3. 03The phishing experiment
  4. 04The format is the failure
  5. 05Completion is not a behaviour
  6. 06What the evidence doesn’t show
  7. 07Training that moves behaviour
© 2026 FUTURE PROOF™
The route. 7 sections, from “The largest training programme on earth” to “Training that moves behaviour”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

At any given moment, somewhere in the world, a huge number of adults are clicking “Next” through a slide about ethics, data protection, workplace safety or suspicious email. Not because they chose to learn this today — because an annual deadline said so. Compliance training in its broad sense is the mandated bundle of ethics, conduct, privacy, safety and security-awareness modules. It is plausibly the largest compulsory education programme in existence.

It is near-universal across large employers, required by regulators and insurers, and refreshed annually. The employee-hours it consumes run to the billions every year. It is also, by a wide margin, the least evaluated education programme of its size. The metric it reports about itself, almost everywhere, is completion.

That would be defensible if the underlying programmes were known to work — if a completed module reliably meant changed behaviour. This article assembles the evidence on that question, and the evidence is uncomfortable. Meta-analyses of ethics instruction find small effects that depend heavily on design (Waples et al., 2009) (Medeiros et al., 2017). The largest randomized field experiment ever run on security training inside a company found its embedded-training arm produced no measurable protection — and by one analysis, harm (Lain et al., 2022). None of this should have been surprising, because the annual massed format violates most of what the learning sciences established decades ago (Cepeda et al., 2006) (Blume et al., 2010). The constructive part is that the same literature is unusually clear about what would work instead.

The largest training programme on earth

Compliance training grew the way it did for reasons that have little to do with how people learn. Regulation demanded proof of diligence; insurers and auditors demanded records. Legal exposure made “we trained everyone” a valuable sentence to be able to say. The product that best satisfies those demands is easy to assign, easy to complete, and easy to document. That is a precise description of the annual module — and an equally precise description of a weak learning intervention. Every incentive in the transaction points at coverage; no party in it is paid for behaviour change.

This is worth stating without cynicism, because the obligations are real. Organizations genuinely must keep data safe, treat people lawfully, and operate without bribery or harassment — and regulators are right to demand evidence of effort. The question this article asks is narrower and more practical: does the standard format discharge those obligations in behaviour, or only in paperwork? That question has an evidence base, and it splits cleanly into two bodies of work. One measures whether the content changes people. The other measures whether the format could ever have worked in the first place (Kirkpatrick & Kirkpatrick, 2006).

The scale claim deserves its own hedge, since this article is about to demand measurement rigour from everyone else. Nobody audits global compliance-training hours. “Billions” is an order-of-magnitude estimate built from near-universal annual mandates across a workforce of hundreds of millions — not a measured statistic. The argument that follows does not depend on the exact figure. It depends on a ratio: the programme consumes effort at roughly this scale. Its outcome evidence would fit in a briefcase — and its best single experiment came back empty-handed (Lain et al., 2022).

What the ethics meta-analyses found

Business-ethics instruction is the corner of the compliance world with the most mature outcome literature — universities and professional schools were evaluating it before corporate programmes existed at scale. Waples, Antes, Murphy, Connelly and Mumford pooled the controlled evaluations and reached a conclusion the field still quotes. The overall impact of business-ethics instruction on ethical perceptions, awareness and decision-making was small — “minimal” was the authors’ own summary word (Waples et al., 2009). The impact also varied widely from programme to programme. That variation was the useful part. Programmes built around focused objectives, active practice and case-based work meaningfully beat broad, lecture-style coverage; design explained far more of the outcome than duration did (Waples et al., 2009).

The follow-up literature sharpened rather than reversed that verdict. Medeiros and colleagues, reviewing the accumulated evaluations of ethics education, found that instruction can produce meaningful gains. The gains concentrate in programmes that are active, professionally situated, and built around the actual dilemmas of the learner’s field — not in generic awareness coverage (Medeiros et al., 2017). Read together, the two syntheses say something more precise than the claim that ethics training does not work. They say it is a design-sensitive intervention deployed, in most organizations, in its least effective design — broad, passive, annual, and identical for everyone (Waples et al., 2009) (Medeiros et al., 2017).

The phishing experiment

Security-awareness training has something the rest of the compliance world lacks: a behavioural outcome that can be measured continuously and objectively. A simulated phishing email either gets clicked or it doesn’t. That made possible the study this literature had been waiting for. Lain, Kostiainen and Capkun ran a randomized experiment across roughly fourteen thousand employees of a partner company over fifteen months — the largest and longest phishing field study published to date. It tested the industry’s standard interventions in real inboxes under real working conditions (Lain et al., 2022).

The number

≈ 14,000 Employees in the largest and longest phishing field experiment published to date — fifteen months, real inboxes, randomized arms (Lain et al., 2022). This is the study the industry’s default product had never faced before.

Two findings matter here. First, simple contextual warnings on suspicious emails helped. Second — the uncomfortable one — embedded training did not reduce later susceptibility. Embedded training is the teachable-moment page shown after an employee falls for a simulated phish. And in the study’s long-run analysis, employees receiving simulated phishing with embedded training went on to fall for phishing more often than colleagues who received no training at all (Lain et al., 2022).

The authors were careful with the mechanism — one candidate explanation is a misplaced sense of being protected by the programme. They were careful with the scope too: one company, one product family, one country. But the industry’s default intervention had just faced the strongest test ever constructed for it, and failed to beat doing nothing (Lain et al., 2022). A third finding pointed forward. Crowdsourced reporting — employees flagging suspicious emails through a one-click button — produced a fast, operationally useful detection signal. The measurable behaviour, in other words, turned out to be more valuable than the training (Lain et al., 2022).

The catch

The teachable moment is an empirical claim, not a truism — and in its strongest test to date it failed to beat doing nothing, with a misplaced sense of being protected as one candidate mechanism. Any embedded-training programme that has never been run against a control is asserting, not measuring.

The result lands harder because the wider awareness literature had predicted it. Bada, Sasse and Nurse reviewed why security-awareness campaigns fail to change behaviour, and catalogued the standard defects years earlier. Campaigns transfer knowledge and assume behaviour will follow; they lean on fear, which fatigues. They are generic where threats are contextual, and one-shot where behaviour change requires sustained, feasible, specific guidance (Bada et al., 2019). Knowing what phishing is has never been the binding constraint. Doing the right thing with a plausible email on a busy Tuesday is — and that is a behaviour, not a fact (Bada et al., 2019).

Annual massed module Spaced boosters + practice 0 3 6 9 12 Months since programme start — illustrative, not measured data © 2026 FUTURE PROOF™
Figure 1. The format argument, drawn to shape: one annual massed module against a spaced-booster programme with practice (dots mark boosters), tracked over twelve months. Illustrative schematic only, drawn after the qualitative shape of the distributed-practice synthesis (Cepeda et al., 2006) and the transfer evidence (Blume et al., 2010) — not measured data from any compliance programme; real trajectories vary with content, population, measure and design, and the vertical axis is deliberately unscaled. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The format is the failure

Here is the deeper problem: even if every module’s content were excellent, the standard delivery format would sabotage it. The annual massed session is very nearly a textbook construction of the condition the memory literature uses as its control group. Cepeda, Pashler, Vul, Wixted and Rohrer synthesised hundreds of experimental comparisons of spaced versus massed practice. The result has anchored the field since: spreading practice across time reliably beats massing it into one session. And the best gap between sessions grows with how long the material must be kept (Cepeda et al., 2006). A behaviour that must hold up for twelve months, trained in a single sitting, is precisely the design this literature would choose to demonstrate forgetting (Cepeda et al., 2006).

The module’s assessment compounds the problem. The end-of-module quiz — multiple choice, retakeable until passed — measures recognition of material shown minutes ago. It does so at the most flattering moment that will ever exist for it. It certifies a peak, then the programme walks away during the decline. Nothing in the format ever samples the quantity the programme exists to change: what the employee does, months later, at the moment of temptation or attack.

Design rule

Measure where the behaviour lives, not where the module ends. A retakeable quiz at the most flattering moment certifies a peak; the programme’s real outcome sits months downstream, so the check has to sit there too — delayed, behavioural, and compared against a baseline.

And even genuine learning would still have to survive the trip back to the job. Blume, Ford, Baldwin and Huang ran the largest pooling of whether trained capability shows up in work behaviour — the meta-analysis of training transfer. Transfer depends on trainee motivation, supervisor and environmental support, and the opportunity to perform what was learned (Blume et al., 2010). The overall relationships were modest, and self-reported transfer flattered the picture. A compulsory module completed alone at a desk supplies none of those conditions: no manager involvement, no practice chance, no follow-up — and a captive audience whose motivation is the deadline (Blume et al., 2010). “Trained” and “changed” are different variables, and the format only ever touches the first.

Completion is not a behaviour

The measurement failure is the one that keeps the others invisible. Kirkpatrick’s four-level framework — reactions, learning, behaviour, results — has organised training evaluation for half a century. Its enduring lesson is that the levels are different questions requiring different evidence (Kirkpatrick & Kirkpatrick, 2006). Compliance dashboards do not typically report level three (behaviour) or level four (results); most do not even report level two — learning measured at a delay. The universal metric, completion, sits beneath level one: it records attendance for a programme whose entire purpose lives at level three (Kirkpatrick & Kirkpatrick, 2006). An organization can therefore run this programme for a decade, hit 100% every year, and know nothing about whether it has ever changed anything.

The standard defence is that behaviour is too hard to measure — and the phishing study quietly removed it. A one-click reporting button produced a continuous, objective, per-employee behavioural metric at trivial cost. The resulting signal proved more useful in practice than the training programme it sat beside (Lain et al., 2022). Behavioural measures of this kind — reporting rates, drill outcomes, decision quality under simulation — are available to any programme that wants them. Organizations measure completion not because behaviour cannot be measured, but because completion is the number the incentive structure asks for (Kirkpatrick & Kirkpatrick, 2006).

The best large-scale illustration of the exposure-versus-structure distinction comes from organizational sociology. Dobbin, Kalev and Kelly analysed three decades of workforce data across hundreds of US firms. They examined which corporate practices actually moved the outcomes those practices targeted. Broadcast training exposure on its own showed little average effect. Structures that assigned responsibility for the outcome — dedicated roles, committees, plans with owners — showed the strongest ones (Dobbin & Kalev, 2006).

The point, for present purposes, is neutral and general. Completing an exposure is not the same event as changing an organizational behaviour. Programmes built as exposure-with-documentation reliably produce documentation (Dobbin & Kalev, 2006). If the outcome you want is behaviour, the unit of both design and measurement has to be behaviour.

What the programme is for, and what it reports Evidence demanded (ordinal) → 100% every year measurement ceiling seldom reported the level compliance needs0 Completion what is reported 1 Reactions did they like it 2 Learning did they learn 3 Behaviour did they change 4 Results did outcomes move © 2026 FUTURE PROOF™
Figure 2. The measurement gap, drawn on the ladder that names it. Kirkpatrick’s four levels ask progressively harder questions of a training programme, and a compliance programme exists to move level three — behaviour, months later, at the moment of temptation. What the dashboard reports is completion: a number that sits beneath level one, reaches 100% every year, and is consistent with a decade of no behavioural change at all. Ordinal, not measured: column height marks how demanding a level’s evidence is, never an effect size. Levels after Kirkpatrick & Kirkpatrick (2006); the cheap level-three metric this article points to — a one-click phishing report button — comes from Lain, Kostiainen & Capkun (2022). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Phishing in organizations: Findings from a large-scale and long-term study. Lain, Kostiainen & Capkun, IEEE Symposium on Security and Privacy, 2022

What the evidence doesn’t show

The uncomfortable reading has boundaries of its own, and they matter as much as the headline.

  • That compliance training “doesn’t work.” The meta-analytic record shows small average effects with strong design moderators — active, case-based, professionally situated instruction produces real gains (Waples et al., 2009) (Medeiros et al., 2017). The null results attach to formats, not to the enterprise.
  • That the phishing result generalizes to all security training. It is one company, one training product family, one country — the study unseats the assumption of benefit and the practice of not measuring; it does not prove all embedded training everywhere is harmful (Lain et al., 2022).
  • That awareness is worthless. Knowledge is necessary — it is simply not sufficient, and campaigns fail when they stop at it (Bada et al., 2019).
  • Long-run cultural effects, in either direction. Whether years of mandatory training shape norms, climate or trust is essentially unmeasured — because behaviour-level measurement is missing from the standard programme in the first place (Kirkpatrick & Kirkpatrick, 2006).
  • That more hours would fix it. Dosage is untested as a remedy, and the moderator evidence points at design features — activity, cases, spacing, context — rather than duration (Waples et al., 2009) (Cepeda et al., 2006).
  • That the legal function is void. Documentation and demonstrable diligence have genuine regulatory value. That is simply a different claim from “this changes behaviour,” and the two should stop being priced as one.

Where the evidence stops

  1. 1That compliance training “doesn’t work.”
  2. 2That the phishing result generalizes to all security training
  3. 3That awareness is worthless
  4. 4Long-run cultural effects, in either direction
  5. 5That more hours would fix it
  6. 6That the legal function is void
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Training that moves behaviour

The same bodies of work that condemn the fire-hose format spell out its replacement with unusual precision.

Space the dose across the year. Replace the annual block with short, recurring sessions whose spacing matches the retention horizon — the single most reliable prescription in the learning sciences (Cepeda et al., 2006). Ten minutes a month is a better-designed programme than two hours a year, at identical cost in employee time.

Practise the behaviour, with feedback — then verify the feedback works. Scenario decisions, realistic drills and case dilemmas are where the ethics evidence concentrates (Waples et al., 2009) (Medeiros et al., 2017). But the phishing experiment is a standing warning: a teachable moment is an empirical claim, not a truism. Instrument it, compare it against controls, and be prepared for the answer (Lain et al., 2022).

Move training to the moment of need. Generic annual coverage fails predictably. Guidance that is contextual, specific and feasible at the point of decision is what the awareness literature says survives contact with behaviour (Bada et al., 2019).

Build the transfer climate. Manager reinforcement, the chance to apply, and follow-up are the strongest levers the transfer evidence offers (Blume et al., 2010). They live outside the module entirely. A programme that budgets nothing for the environment is betting against its own meta-analysis.

Measure behaviour, and give the outcome an owner. Report level-three measures — reporting rates, drill performance, decision quality, incident response — not completion percentages (Kirkpatrick & Kirkpatrick, 2006). And assign responsibility for the outcome to named structures rather than to broadcast exposure. That is the strongest pattern in the large-scale organizational record (Dobbin & Kalev, 2006).

Applied at Future Proof

How Future Proof™ applies this.

The evidence says compliance outcomes are behaviours, and behaviours are built by spacing, practice, context and measurement — not by an annual module. Future Proof runs mandated topics the way the literature specifies: content is decomposed into short micro-boosters scheduled across the year against a forgetting model, so the dose lands where the decay is. Practice takes the form of scenario decisions with immediate, process-level feedback rather than recognition quizzes. And the analytics report behaviour — drill performance over time, decision quality, response latency — alongside the completion record the auditors still need. Because every intervention is instrumented, the teachable moment is tested rather than assumed, and formats that fail their own data get redesigned instead of re-assigned.

See the platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 2006Cepeda
  • 2006Kirkpatrick
  • 2006Dobbin
  • 2009Waples
  • 2010Blume
  • 2017Medeiros
  • 2019Bada
  • 2022Lain
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 2006–2022, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Waples, E.P., Antes, A.L., Murphy, S.T., Connelly, S., & Mumford, M.D. (2009). A meta-analytic investigation of business ethics instruction. Journal of Business Ethics 87: 133–151. PDF
  2. Medeiros, K.E., et al. (2017). What is and what can be: The scope and possibilities of ethics education. Ethics & Behavior 27(5): 351–384. PDF
  3. Lain, D., Kostiainen, K., & Capkun, S. (2022). Phishing in organizations: Findings from a large-scale and long-term study. IEEE Symposium on Security and Privacy: 842–859. PDF
  4. Bada, M., Sasse, A.M., & Nurse, J.R.C. (2019). Cyber security awareness campaigns: Why do they fail to change behaviour? arXiv:1901.02672. PDF
  5. Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132(3): 354–380. DOI
  6. Blume, B.D., Ford, J.K., Baldwin, T.T., & Huang, J.L. (2010). Transfer of training: A meta-analytic review. Journal of Management 36(4): 1065–1105. PDF
  7. Kirkpatrick, D.L., & Kirkpatrick, J.D. (2006). Evaluating Training Programs: The Four Levels (3rd ed.). Berrett-Koehler. PDF
  8. Dobbin, F., & Kalev, A. (2006, with Kelly). Best practices or best guesses? Assessing the efficacy of corporate affirmative action and diversity policies. American Sociological Review 71(4): 589–617. PDF
Try the AI engine

Compliance that changes behaviour.

Book a 20-minute demo. We’ll show you spaced micro-boosters scheduled against forgetting, scenario practice with feedback, and behavioural measurement your auditors and your security team can both live with.

8 citations Reviewed August 2026 Open peer review welcomed