Leadership training: better than its reputation.
Leadership development is the largest single line in many corporate learning budgets, and “leadership training doesn’t work” may be the most quoted sentence in the industry. Four decades of meta-analysis say something more interesting: it works — on average, clearly — and a short list of design choices decides how much. What the evidence says, moderator by moderator.
The finding: Leadership training has one of the strongest evidence bases in corporate learning. The modern anchor — the Lacerenza et al. meta-analysis — found positive effects across reactions, learning, transfer, and organizational results, with learning effects on the order of d ≈ 0.7 and transfer effects around d ≈ 0.8, both approximate. Earlier syntheses reaching back to 1986 agree on the direction.
The mechanism: The average conceals the story: design moderators decide the outcome. Programs built on a needs analysis, feedback, multiple spaced sessions, and practice-based methods substantially outperform one-shot lectures, and self-administered programs underperform. The winning list is the learning-science checklist — spacing, practice, feedback — applied to leadership.
The product: Future Proof™ builds those moderators into the delivery layer — spaced practice schedules, scenario practice with feedback from the AI Tutor, and analytics that measure behavior change rather than satisfaction — so leadership curricula run on the design features the meta-analyses reward.
In this article
- 01The most cynical line item in the budget
- 02Four decades of the same answer
- 03The modern anchor: Lacerenza and colleagues
- 04The moderators are the finding
- 05A worked example: behavior modeling training
- 06Transfer: where good programs go to die
- 07Programs are episodes. Development is a process.
- 08The honest caveat: variance is the norm
- 09What the evidence doesn’t show
- 10What this means for practice: buy features, not brands
Corporate leadership development has a strange reputation problem. It eats the largest share of many learning budgets — by most industry accounts, more than any other single category of training. Yet the sentence “leadership training doesn’t work” circulates so freely that it has become a genre of business commentary in its own right. Executives repeat it after off-sites that changed nothing. Journalists repeat it about programs that were never designed to be evaluated. The spend and the cynicism rise together, and neither side seems to feel obliged to check the literature.
The literature, it turns out, has an answer, and it has been giving roughly the same one for four decades. Managerial and leadership training works on average. Meta-analysis after meta-analysis finds genuine effects on knowledge, behavior, and results. And the effects vary hugely with how programs are designed. That second clause is the interesting one.
The moderators that separate strong programs from weak ones are not mysterious and not secret. They are the standard checklist of learning science — spacing, practice, feedback, fit to diagnosed need — applied to the most costly training in the building. This article walks the evidence in order. First the early meta-analyses, then the modern anchor study, then the design moderators that carry the story — and last, the transfer literature that explains where even well-designed programs go to die.
The most cynical line item in the budget
Start with the shape of the market, because it explains the shape of the cynicism. Leadership development is the rare corporate purchase made at huge scale yet almost never checked beyond a satisfaction survey. Programs are picked on brand, faculty fame, or the pull of a keynote. Their outcomes are measured — when they are measured at all — by asking people whether they enjoyed themselves. In that vacuum, both extreme positions flourish. Vendors claim transformation; critics note that transformation rarely survives contact with Monday morning; and “leadership training doesn’t work” hardens from complaint into common knowledge.
The complaint deserves to be taken seriously, because parts of it are true. Plenty of programs are inspirational theater. Ropes courses, personality games, and two-day retreats with no follow-up have exactly the staying power you would expect. But the general claim — that leadership training as a category does not work — is a claim about evidence. And it happens to be one of the better-tested claims in applied psychology, because leadership training is among the most meta-analyzed programs in the field’s literature. The meta-analyses, spanning four decades of studies, keep returning the same two-part verdict: positive on average, wildly variable by design.
Four decades of the same answer
The record of that verdict is old enough to embarrass the cynics. Burke and Day published the first major pooled analysis in 1986, combining decades of managerial-training studies. They found training to be, on average, at least moderately effective across content areas and outcome measures (Burke & Day, 1986). That paper is old enough to have trained the trainers of today’s trainers. Its central finding has never been overturned — only refined.
Two later syntheses did the refining. Collins and Holton pooled studies of managerial leadership development from 1982 to 2001. They found real gains in both knowledge and performance — with the caveat that effects varied widely across programs, outcome types, and study designs (Collins & Holton, 2004). Avolio and colleagues then took the broadest cut: a review of roughly 200 leadership intervention studies, true experiments and their close cousins. It found a positive causal impact of leadership interventions on work outcomes — real, but modest on average (Avolio et al., 2009). Note the word “causal”: this literature contains genuine experiments and field trials, not just before-and-after surveys of the converted.
Read together, the early syntheses establish two facts the cynical consensus cannot digest. First, the direction is consistent: across eras, methods, and outcome types, the average effect of managerial and leadership training is positive. Second, the variance is enormous — some programs move nothing, some move a great deal. That is precisely the pattern you would expect if design quality, rather than category futility, were the active ingredient. What the early work could not do was test that idea directly. That took until 2017.
The modern anchor: Lacerenza and colleagues
The test arrived as the anchor study of the modern literature: Lacerenza, Reyes, Marlow, Joseph, and Salas’s meta-analysis of leadership training programs, published in the Journal of Applied Psychology (Lacerenza et al., 2017). It sorted outcomes on the classic four-level ladder: reactions (did people like it), learning (do they know or can they do more), transfer (did behavior change on the job), and results (did organizational outcomes move). And it asked both the average question and, crucially, the design question.
The averages alone would have justified the paper. Leadership training produced positive effects at every level of the ladder. Learning effects were substantial — on the order of d ≈ 0.7, approximately. Transfer effects — measured at the level where training is supposed to go to die — were roughly as large or larger, around d ≈ 0.8, again approximate. Results-level effects on organizational outcomes were positive and meaningful as well, resting on a smaller base of studies (Lacerenza et al., 2017).
Taken at face value, the estimates say trained leaders behaved differently at work afterward — and their firms could measure the difference. For a category routinely described as a write-off, that is a genuine surprise. The effects are moderate to large, by the conventions of applied psychology.
d ≈ 0.8 The approximate transfer-level effect of leadership training — behavior change measured on the job, roughly as large as or larger than the learning effect behind it (Lacerenza et al., 2017).
But a meta-analytic average is an abstraction over hundreds of very different programs. The paper’s more important contribution was to break the abstraction apart. Which design features separated the programs that produced those effects from the programs that produced nothing? The answer is the closest thing corporate learning has to a buying guide — and it is the story the rest of this article is about.
The moderators are the finding
Here is the moderator pattern, feature by feature. Programs that began with a needs analysis — a diagnosis of what these leaders in this organization actually need to learn — beat programs that started from a generic curriculum. Programs that included feedback beat programs that only presented content. Programs delivered as multiple sessions spaced over time beat single massed events.
Programs built on practice — cases, role-plays, simulations, application on the job — beat designs made of lectures and information. Programs mixing hard skills with people skills beat one-track offerings. And one delivery mode reliably fell short. Programs learners work through alone, with no instructor, produced weaker outcomes than instructor-led ones (Lacerenza et al., 2017).
Buy features, not brands: needs analysis first, feedback included, sessions spaced over time, practice-based methods, instructor-led delivery. Two programs teaching identical content can sit at opposite ends of the effectiveness distribution on these alone.
Look at that list again, because its structure is the point. Spacing over massing. Practice over passive exposure. Feedback over none. Fit to diagnosed need over generic content.
These are not discoveries about leadership as such. They are the general findings of a century of learning science, showing up on schedule in the most costly corner of corporate training. The moderator table of a leadership meta-analysis reads like the contents page of a textbook on how people learn. So the difference between a program that works and a program that doesn’t is rarely the model of leadership being taught. It is whether the delivery respects how humans build durable skills.
It also means the market buys on the wrong variable. Programs are marketed on content and faculty — the brand of the leadership framework, the fame of the school. The evidence says outcomes track design features that appear nowhere in a brochure. The number and spacing of sessions; the ratio of practice to presentation; the presence of structured feedback; the diagnostic work done before day one. Two programs teaching identical content can sit at opposite ends of the effect range on delivery design alone. That is what the figure below shows, schematically: the same program, with and without each feature, is effectively two different products.
A worked example: behavior modeling training
If the moderator list is the theory, behavior modeling training is the demonstration. It is the leadership-adjacent method that has always shipped with the winning design features built in. A behavior modeling course teaches a skill — say, giving corrective feedback without triggering defensiveness — in a fixed cycle. State explicit behavior rules. Show a model performing them. Then cycle every trainee through practice and feedback until the behavior is fluent, with structured support for using it back on the job.
Taylor, Russ-Eft, and Chan meta-analyzed the method across the accumulated study base. They found what the moderator story predicts: solid effects on knowledge and skill, and — the finding that matters — durable changes in on-the-job behavior. The effects were largest under three conditions. Training included retention aids, such as rule codes trainees could carry into the workplace. Trainees practiced with scenarios drawn from their own jobs. And the post-training environment supported transfer through rewards and consequences (Taylor, Russ-Eft & Chan, 2005).
Effects on job behavior held up over time in a way that training that merely delivers content rarely matches. The lesson generalizes. When a leadership skill is taught the way learning science says skills are built — model, practice, feedback, spaced use, support in the environment — it behaves like any other trainable skill. The mystique of unteachable leadership evaporates at exactly the point where the teaching becomes rigorous.
Transfer: where good programs go to die
Even a well-built program, however, controls only half the outcome. None of this survives contact with the workplace unless the workplace cooperates, and the transfer literature is blunt about how often it doesn’t. Blume, Ford, Baldwin, and Huang meta-analyzed transfer of training across occupations. Whether trained skills actually appear on the job depends heavily on conditions the classroom never sees. Those conditions: support from supervisors and peers, the chance to perform the new skills soon after training, and the general climate of the post-training environment (Blume et al., 2010). Trainee traits — cognitive ability, conscientiousness, motivation — matter too, but the levers in the environment are the ones a company actually controls.
A course only controls half the outcome. Whether trained skills appear on the job tracks supervisor and peer support, early opportunity to perform, and transfer climate (Blume et al., 2010) — so treat the participant’s manager, and the first months after the last session, as part of the curriculum.
This is the missing half of every “the training didn’t work” story. A leadership program can produce genuine learning that then decays untouched. The participant returns to a manager who never asks about it, a calendar with no room to practice it, and an incentive system indifferent to it. The Lacerenza moderators govern what happens inside the program; the Blume moderators govern what happens after it; and a serious development effort has to be designed across both. Transfer is not a property of a course; it is a property of a course plus an environment. That is why the strongest programs treat the participant’s manager, and the first months after the last session, as part of the curriculum rather than as the audience.
Leadership training design, delivery, and implementation.Lacerenza et al., Journal of Applied Psychology, 2017 — the moderators are the finding
Programs are episodes. Development is a process.
Zoom out one more level and the program itself starts to look small. Day, Fleenor, Atwater, Sturm, and McKee reviewed 25 years of leader-development research and theory. Their case: leadership develops over the long haul, through sequences of experience, feedback, and reflection unfolding across years (Day et al., 2014). Formal programs, on that reading, are episodes within that process rather than the process itself. The distinction sounds philosophical and is intensely practical. If development is a process, then “did the program work?” matters less than “what happens between programs?” — who gets stretch assignments, who receives feedback dense enough to learn from, whose reflection is supported rather than left to chance.
The process view also explains why spacing shows up so strongly in the program-level data. A leadership curriculum spread over months, with real work between sessions, is not merely a better-scheduled course. It is a small-scale model of how leader development actually works: cycles of teaching, attempt, feedback, and consolidation. The design moderator and the development theory are the same claim at two zoom levels — a reassuring thing for a literature to produce by accident.
The honest caveat: variance is the norm
If the evidence is this consistent, why does the cynicism survive? Partly because the variance is real — and anyone selling certainty in either direction should be made to read Powell and Yalcin. Their meta-analysis covered managerial training from 1952 to 2002. Average effects varied widely by outcome measure and by era — respectable in places, unimpressive in others, with no triumphant upward trend across the half-century (Powell & Yalcin, 2010). Their reading is a useful astringent: the average program, averaged over fifty years of practice, was no revolution.
But variance cuts in a particular direction once you know what the moderators are. A wide spread with a positive mean is exactly what a design-sensitive intervention looks like from a distance. The left tail is full of massed, lecture-based, feedback-free events; the right tail is full of programs that would satisfy the Lacerenza checklist. The cynics are not hallucinating — they are accurately describing the left tail, which may well be where most of the budget goes. The error is generalizing from the tail to the category. “Leadership training doesn’t work” and “leadership training works” are both false as stated; the defensible sentence is that leadership training works when it is designed right, and the design features are known.
What the evidence doesn’t show
The case above is strong, and it has edges. Six limits are worth keeping in view before anyone converts a meta-analysis into a procurement memo:
- Most outcomes are ratings, not business results. Much of the transfer evidence rests on behavior ratings by supervisors, peers, or trainees themselves — better than satisfaction scores, still softer than revenue, retention, or safety. Results-level effects are positive but rest on far fewer studies.
- Publication bias flatters everything. Meta-analyses inherit the file drawer: null program evaluations are less likely to be written up at all, so the true average effects are plausibly smaller than the published ones.
- “Leadership” is not one construct. The primary studies mix supervisory skills, charisma, strategic thinking, and safety leadership under a single label; pooling heterogeneous constructs makes the average an average over apples and engines.
- Long-run durability is thin. Follow-ups measured in months are common and follow-ups in years are rare, so how much trained behavior survives a promotion, a reorganization, or a new boss is largely unmeasured.
- Participants are rarely random. Organizations send the already-promising to leadership programs, and volunteers differ from conscripts; self-selection can flatter estimates of what training would do for everyone else.
- Moderator evidence is correlational across studies. Spaced versus massed is mostly a comparison between different studies, not a randomized head-to-head within one, so confounds between design choices and organizational quality cannot be fully excluded.
Where the evidence stops
- 1Most outcomes are ratings, not business results
- 2Publication bias flatters everything
- 3“Leadership” is not one construct
- 4Long-run durability is thin
- 5Participants are rarely random
- 6Moderator evidence is correlational across studies
What this means for practice: buy features, not brands
Turning this literature into practice is almost embarrassingly direct: buying should read like the moderator table. Before content, before faculty, before brand, put a leadership program through the design questions the evidence rewards. Was a needs analysis done — on this organization, not the industry in general — and is delivery spaced over multiple sessions with workplace application between them? What fraction of contact time is practice with feedback rather than presentation? Does the design include the participant’s manager and the post-program environment, or does it end at the classroom door (Blume et al., 2010)? And is anything measured beyond satisfaction — ideally behavior, rated by the people who work with the leader?
A checklist like that filters the market brutally, because the features it demands are costly in exactly the way inspirational events are not. Spacing costs calendar. Practice costs facilitation. Feedback costs measurement. Needs analysis costs diagnostic work before the invoice. But the meta-analytic gap between programs with and without those features is the whole difference between an expense and an investment (Lacerenza et al., 2017). And unlike the brand of the leadership model, every item on the list can be checked before signing.
The deeper shift is to treat leadership development as an operating system rather than an event stream. Diagnosed needs feed spaced curricula. Every session cycles through practice and feedback. Managers are enlisted as transfer infrastructure, and behavior is measured on the job as a matter of course (Day et al., 2014). Nothing in that list is speculative. Every clause is a moderator with a meta-analysis behind it — and together they turn the most cynical line item in the budget into one of the better-evidenced things an organization can buy.
How Future Proof™ applies this.
The moderators that decide whether leadership training works are delivery infrastructure — which is what Future Proof is. Curricula run as spaced schedules rather than events, so practice lands where memory needs it instead of where the calendar allowed it. The AI Tutor turns scenarios into deliberate practice with immediate feedback, the knowledge map ties content to diagnosed gaps rather than generic competency lists, and assessments measure behavior change over time — with manager-visible analytics standing in for the transfer support the evidence demands.
See the platform →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above.
The evidence, by year
- 1986Burke
- 2004Collins
- 2005Taylor
- 2009Avolio
- 2010Blume
- 2010Powell
- 2014Day
- 2017Lacerenza
- Lacerenza, C.N., Reyes, D.L., Marlow, S.L., Joseph, D.L., & Salas, E. (2017). Leadership training design, delivery, and implementation: A meta-analysis. Journal of Applied Psychology 102(12): 1686–1718. DOI
- Burke, M.J., & Day, R.R. (1986). A cumulative study of the effectiveness of managerial training. Journal of Applied Psychology 71(2): 232–245. PDF
- Collins, D.B., & Holton, E.F. (2004). The effectiveness of managerial leadership development programs: A meta-analysis of studies from 1982 to 2001. Human Resource Development Quarterly 15(2): 217–248. PDF
- Avolio, B.J., Reichard, R.J., Hannah, S.T., Walumbwa, F.O., & Chan, A. (2009). A meta-analytic review of leadership impact research: Experimental and quasi-experimental studies. The Leadership Quarterly 20(5): 764–784. PDF
- Taylor, P.J., Russ-Eft, D.F., & Chan, D.W.L. (2005). A meta-analytic review of behavior modeling training. Journal of Applied Psychology 90(4): 692–709. PDF
- Blume, B.D., Ford, J.K., Baldwin, T.T., & Huang, J.L. (2010). Transfer of training: A meta-analytic review. Journal of Management 36(4): 1065–1105. DOI
- Day, D.V., Fleenor, J.W., Atwater, L.E., Sturm, R.E., & McKee, R.A. (2014). Advances in leader and leadership development: A review of 25 years of research and theory. The Leadership Quarterly 25(1): 63–82. PDF
- Powell, K.S., & Yalcin, S. (2010). Managerial training effectiveness: A meta-analysis 1952–2002. Personnel Review 39(2): 227–241. PDF
Run leadership development on the evidence.
Book a 20-minute demo. We’ll show you how spaced schedules, scenario practice with feedback, and behavior-level measurement turn the meta-analytic checklist into your default delivery — for every leadership cohort.