Reskilling at scale: what the trials show.
Governments have run the world’s largest reskilling experiments for decades, and the results are meta-analyzed: training works — with a delay, unevenly, and dramatically better in some designs than others. What the evidence says about which designs, and what corporate reskilling should copy from the programs that actually moved earnings.
The finding: Across hundreds of evaluations, active labor-market training programs show small or even negative effects in the first year and meaningful positive effects two to three years out — training is an investment with a J-curve, not a quick fix. The star performers are sector-based programs that train for specific, verified employer demand: randomized trials of these show large, durable earnings gains that ordinary programs never approach.
The mechanism: Generic training builds supply and hopes demand appears; sectoral designs verify demand first, train to the actual hiring bar, and connect graduates to employers. The demand linkage — not the classroom — is the differentiator.
The product: Future Proof’s reskilling tooling is built demand-first: target roles defined from real internal vacancies and role maps, gaps diagnosed against the destination bar, and progress measured as placement-readiness rather than course completion.
In this article
- 01The J-curve: training works late
- 02The sectoral exception that should be the rule
- 03Why the counterfactual discipline matters
- 04Translating to the corporate case
- 05The automation frame, kept honest
- 06What the learning science adds
- 07What the evidence doesn’t show
- 08What this means for practice
Every automation headline ends in the same corporate vow: we will reskill our people. The vow is sincere and the budgets are real. The follow-through is usually a content library, an enrollment target, and a hope. That pattern is not chosen because anyone judged it best. It is chosen because nobody asked the evidence what works, at what pace, for whom.
Reskilling is the rare topic where corporate strategy could consult a huge body of controlled evidence — and almost never does. For fifty years, governments facing job losses — recessions, closing industries, automation — have funded retraining at population scale. Economists have tested those programs with the field’s full toolkit: natural experiments, randomized trials, and meta-analyses spanning continents. Corporate talk proceeds largely as if none of this exists. That is a pity, because the evidence answers the questions executives actually ask: does retraining work, how long does it take, and what separates the programs that lift earnings from the ones that merely fill time?
The J-curve: training works late
For most of the late twentieth century, the informed answer to “does government retraining work?” was a shrug trending negative. That skepticism was earned honestly, from early programs and early methods alike (Heckman et al., 1999). What changed the answer was volume and rigor arriving together: hundreds of evaluations, more and more of them randomized, spanning enough countries and cycles to support real pooling.
The big meta-analysis of recent active labor-market programs — over two hundred studies from around the world — turned the field’s gloom into something more precise. Averaged naively, training programs look weak: near-zero or negative employment effects in the first year. Followed further, the picture inverts. Effects turn positive and grow through years two and three, and training is among the program types whose medium-run effects are strongest. The result is a J-curve: participants first fall behind (they are in training, not searching), then durably overtake comparison groups (Card, Kluve & Weber, 2018). Earlier reviews had flagged the same timing, plus the moderators — the factors that move the effect — that persist: programs do better in recessions, better for women and the long-term unemployed, and better when content connects to real labor demand (Kluve, 2010).
The J-curve alone should reshape how companies govern reskilling. Internal programs are routinely judged on exactly the horizon where the evidence predicts nothing: the quarter after the course ends. A program evaluated at ninety days is being graded during its dip. The meta-analytic message: honest measurement windows run eighteen to thirty-six months, and programs killed early are killed on noise (Card et al., 2018). The older skeptical tradition earned its doubts against programs judged — and designed — without that patience (Heckman, LaLonde & Smith, 1999).
A reskilling initiative evaluated at ninety days is being graded during its dip. The honest measurement windows sit at eighteen to thirty-six months — which means a program killed at the quarter mark is killed on noise, and a governance cadence built on quarterly reviews will bury designs the evidence says were working.
The sectoral exception that should be the rule
If the meta-analysis established that training can work, the next question is what separates the programs that do. Here the literature has something close to a solved case — a design whose randomized results sit so far above the field’s averages that its ingredients deserve line-by-line attention. Averages hide the design story, and the design story is the actionable part.
A family of U.S. programs — sector-based training — departs from the generic model on every axis. They pick a target sector with verified hiring demand, screen entrants for readiness, and train to the sector’s actual skill bar in weeks-to-months of intensive work. Then they hand graduates to employer partners who helped define the curriculum. Randomized trials of the flagship programs found earnings gains of a size the generic literature never sees — on the order of 15–30%, sustained years after training. The gains flowed through access to better, higher-wage jobs in the targeted sectors, not merely more hours of work (Katz, Roth, Hendra & Schaberg, 2022). Long-run follow-ups of the strongest sites show the effects persisting toward a decade (Roder & Elliott, 2019).
15–30% Sustained earnings gains in randomized trials of sector-based training — demand-verified, employer-linked, readiness-screened — in a field where moving earnings a few percent counts as success (Katz, Roth, Hendra & Schaberg, 2022).
The scale of those numbers deserves a pause. Labor economics is a field where a program that moves earnings a few percent counts as a success, and most programs move nothing. The sectoral trials’ sustained double-digit gains are, by the field’s standards, extraordinary — the kind of result that gets replicated skeptically across sites and years precisely because nobody believes it at first. It has been, repeatedly (Schaberg, 2017).
The trials’ internal contrasts point to the active ingredient. People who finished training but were not placed through the demand-side machinery gained far less; sites with weaker employer links produced weaker effects. And the programs’ own theory of change — train for jobs that verifiably exist, to the standard employers verifiably require — is the through-line the evaluations keep confirming (Katz et al., 2022). The lesson travels almost embarrassingly well: reskilling fails as a supply-side ritual and works as a demand-matched pipeline.
Why the counterfactual discipline matters
The word “evaluation” does heavy lifting throughout this article. Its content is the discipline most corporate programs skip: a comparison group. Without one — without a counterfactual, an estimate of what would have happened anyway — reskilling outcomes cannot be read, in exactly the ways our measurement cluster warns.
People self-select: the driven apply, so raw before-after gains blend the program’s effect with the applicants’ own path. The economy moves: a cohort reskilled into a boom looks brilliant no matter what. And regression to the mean guarantees that programs recruiting recently laid-off workers will show “recovery” no training caused. Labor economists spent two decades learning these lessons the hard way. That is why the modern literature runs on randomization and careful quasi-experiments (Heckman, LaLonde & Smith, 1999) — and why its numbers deserve the trust that corporate case studies built on testimonials do not.
Inside a firm, the counterfactual is cheaper than it sounds. Oversubscribed programs create natural comparison groups from waitlists. Phased rollouts create them from scheduling. And matched non-participants — same origin role, tenure band, and baseline diagnostics — give a workable benchmark where randomization is impossible. The point is not academic purity; it is decision quality at scale-up time. A company about to scale a reskilling model on the strength of its pilot should know whether the pilot’s gains belonged to the program or to the people who signed up first.
Translating to the corporate case
A fair objection: public programs serve jobseekers moving between employers, while corporate reskilling moves people inside one firm — do the lessons transfer? The structural differences are real. Mapped one by one, though, most cut in the corporate program’s favor — provided the design lessons travel with the money. The demand check that sectoral programs work hardest to get — are there really jobs? — is, inside a firm, a query against workforce plans. The destination roles exist, their headcount is budgeted, and their skill bar can be read off the role map rather than guessed. The placement machinery is internal too: no outside employer must be persuaded to interview graduates.
The corporate failure modes are different in matching ways. Programs get aimed at wished-for goals, not budgeted ones (“everyone learns AI”). Course content gets copied from generic vendors, not derived from the target bar. People get recruited by enthusiasm, not diagnosed fit. Success gets counted as courses finished, not people moved into new roles. Each is the generic-program mistake, brought indoors.
The screening finding needs careful translation. Sectoral programs screen for baseline readiness, and their effects partly depend on it. In public policy that raises hard fairness questions; inside a firm it converts cleanly into diagnosis and preparation. Candidates who lack the skills the target role assumes are not turned away. They are routed through the foundation layer first, with the diagnostic data our knowledge-mapping articles describe deciding who needs which on-ramp. A program with honest gap measurement can be pickier about order and, at the same time, more open about who eventually gets in than any outside program — the advantage of owning both the training and the time.
The automation frame, kept honest
Corporate reskilling talk leans on automation projections — millions of roles transformed, timelines attached. The economics counsels a calmer read. The occupational-change literature stresses that automation typically reshapes jobs at the task level rather than deleting them outright: routine parts leave, judgment and people-facing parts grow, and the net picture is churn within roles more than extinction of them (Autor, 2015). For reskilling design, the task-level frame is a gift. It turns an unanswerable question (“which jobs will disappear?”) into a mappable one (“which tasks in our roles are shifting, and what do the growing tasks require?”). That is a question a role map answers and a headline never will.
It also right-sizes the effort. Task-level churn mostly demands the delta training of our skill-half-life review — continuous, modest, targeted. Full reskilling is reserved for the genuine cases where a role’s foundation layer is what changes. Companies that read the automation literature as “reskill everyone, urgently” build the generic supply-side programs the trial evidence buries. Companies that read it at task grain build pipelines sized to actual movement — cheaper, and the only version the evidence endorses.
What the learning science adds
Everything so far concerns program design — demand, screening, placement, patience. The trials say little about teaching method, because the studies treat the classroom as a black box. The box’s contents are this library’s home turf, and for reskilling the contents matter more than usual. The economics grades programs as containers; the learning science says what should be inside them.
Adult reskillers are the group the core findings bite hardest. Prior knowledge is uneven, so diagnosis-first beats marching through a syllabus. Time is scarce, so the efficient tools — spacing, retrieval, worked examples — are not extras; they are what makes the plan fit the hours. Confidence is fragile, so the error-friendly designs of our pretesting and productive-failure reviews also help keep people enrolled, and the low-stakes testing design of our test-anxiety review keeps measurement from driving people out. And because the goal is a job, not a certificate, practice should be simulation-shaped — scenarios from the target role — plus the transfer supports that, per our compliance-training review, decide whether what was trained survives contact with the job.
One more crossover matters at the portfolio level: the make-or-buy math. Hiring a scarce skill from outside prices at market premium plus onboarding time plus failure risk. Reskilling from inside prices at program cost plus the J-curve’s patience. The trials’ effect sizes let you run that math for the first time. And the math increasingly favors building — especially where the destination roles share foundations with the origin roles, which is exactly what a knowledge map can measure in advance (Katz et al., 2022).
Reskilling fails as a supply-side ritual and works as a demand-matched pipeline.The design lesson of the sectoral-training trials — Katz et al. (2022), Roder & Elliott (2019).
What the evidence doesn’t show
- It doesn’t show all training pays. The averages include plenty of programs that never beat their comparison groups; the literature’s gift is discriminating the designs, not blessing the category (Card et al., 2018).
- It doesn’t map one-to-one onto incumbent workers. The trials mostly study unemployed and low-income jobseekers; corporate reskilling of employed workers borrows the design logic on plausibility, not on a parallel trial base — an honest gap worth naming.
- Selection is doing some work in every program. Even randomized designs estimate effects for people who applied; scaling a program changes who walks through the door, and effects with it (Heckman et al., 1999).
- Certificates are not outcomes. The trials measure earnings and placement; any internal program measuring completions is measuring the input and calling it the result — the smile-sheet error at program scale.
Where the evidence stops
- 1It doesn’t show all training pays
- 2It doesn’t map one-to-one onto incumbent workers
- 3Selection is doing some work in every program
- 4Certificates are not outcomes
What this means for practice
The playbook writes itself from the trials, and none of it requires public-sector scale. Start every reskilling effort at the demand end. Name the destination roles; check their headcount is real and budgeted; pull the actual skill bar from the role map and from the people now succeeding in those roles. That is the sectoral programs’ curriculum method, run against internal truth.
Then diagnose the candidate pool against that bar and let the gaps define the program: on-ramps for missing foundations, intense, focused training for the target layer, scenario practice for the judgment parts. Publish the intended transition rate before launch. That number — not sign-ups, not completions, not satisfaction — is what the trial literature would grade. Naming it in advance keeps the program honest when the dip arrives.
Design the experience for working adults, not full-time students: intense but bounded blocks, spaced practice instead of marathon sessions, and visible early wins against the target bar. That is drop-out risk management — the concern the public programs’ drop-out data made first-class. Govern with the J-curve in the deck. Set measurement windows the evidence respects: readiness milestones during training, transition rates at six to twelve months, on-the-job results beyond. And pre-commit leaders to the dip, because the alternative is the documented pattern — programs cut at the moment their data says least.
Track the reskilled group against matched peers. That honors the counterfactual discipline this literature is built on, and it builds the internal evidence that makes the next program’s case. Reskilling’s reputation problem was never that retraining fails; it is that patience and demand links were treated as optional extras rather than the mechanism itself. The trials priced both, at population scale, with control groups. They are not features of the program. They are the program.
How Future Proof™ applies this: demand-first reskilling.
Reskilling in the platform begins where the successful trials begin: at the destination. Target roles carry explicit skill bars from the knowledge map; candidates are diagnosed against them, routed through on-ramps where foundations are missing, and trained with the efficiency stack — spacing, retrieval, scenarios — the compressed timeline demands. Dashboards report readiness against the destination bar and transitions achieved, not courses consumed, and cohort outcomes are tracked on the multi-year horizon the J-curve requires. Supply-side ritual is the one design the system won’t let you build. That constraint is the product’s opinion, and the evidence is on its side.
See reskilling pipelines →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
The evidence, by year
- 1999Heckman
- 2003Arthur
- 2010Kluve
- 2015Autor
- 2015Barnow
- 2017Schaberg
- 2018Card
- 2019Roder
- 2022Katz
- Card, D., Kluve, J., & Weber, A. (2018). What works? A meta-analysis of recent active labor market program evaluations. Journal of the European Economic Association 16(3): 894–931. PDF
- Kluve, J. (2010). The effectiveness of European active labor market programs. Labour Economics 17(6): 904–918. PDF
- Heckman, J.J., LaLonde, R.J., & Smith, J.A. (1999). The economics and econometrics of active labor market programs. In Handbook of Labor Economics, Vol. 3A (Ashenfelter & Card, eds.): 1865–2097. PDF
- Katz, L.F., Roth, J., Hendra, R., & Schaberg, K. (2022). Why do sectoral employment programs work? Lessons from WorkAdvance. Journal of Labor Economics 40(S1): S249–S291. PDF
- Roder, A., & Elliott, M. (2019). Nine Year Gains: Project QUEST’s Continuing Impact. Economic Mobility Corporation. PDF
- Autor, D.H. (2015). Why are there still so many jobs? The history and future of workplace automation. Journal of Economic Perspectives 29(3): 3–30. PDF
- Schaberg, K. (2017). Can Sector Strategies Promote Longer-Term Effects? Three-Year Impacts from the WorkAdvance Demonstration. MDRC. PDF
- Arthur, W., Bennett, W., Edens, P.S., & Bell, S.T. (2003). Effectiveness of training in organizations: A meta-analysis of design and evaluation features. Journal of Applied Psychology 88(2): 234–245. DOI
- Barnow, B.S., & Smith, J. (2015). Employment and training programs. In Economics of Means-Tested Transfer Programs in the United States, Vol. 2 (Moffitt, ed.): 127–234. PDF
Reskill toward verified demand.
Book a 20-minute demo. Bring two destination roles — we’ll show you the skill bars, the gap diagnosis, and the pipeline view that measures transitions, not completions.