Does coaching work? Finally, an answer.
Coaching went from executive perk to a multi-billion-dollar industry before anyone seriously measured whether it worked — an inversion of the usual order that lasted decades. The meta-analyses have now reported. The answer is yes, moderately, with conditions that most corporate programs quietly ignore.
The finding: The meta-analyses converge: workplace coaching produces real, positive, moderate effects — roughly g ≈ 0.4 to 0.7 depending on the outcome domain, largest for goal-directed self-regulation. Internal coaches perform at least as well as external ones, delivery format matters less than expected, and the classic longitudinal study found coached managers improved on multisource ratings — modestly.
The mechanism: Coaching works mainly through goal-directed self-regulation and the quality of the working alliance — specific goals, monitored behaviour, structured follow-through, a coachee who wants to be there. It does not appear to work through the coach’s charisma or credentials. And the strongest adjacent literature shows effects shrink as programs scale, because dosage and quality dilute.
The product: Future Proof™ applies the mechanisms the evidence supports — goal-setting loops, spaced follow-through checks, measured behaviour change — through AI coaching at a dosage that survives scale, with human coaches reserved for the judgment-heavy work where they earn their cost.
In this article
- 01The boom that skipped the measurement
- 02The first real answer
- 03The workplace meta-analysis and its moderators
- 04The honest anchor
- 05What actually drives outcomes
- 06The ROI trap
- 07The scale lesson from teacher coaching
- 08What the evidence doesn’t show
- 09Coaching by the evidence
Almost every workplace program follows the same life cycle: a plausible idea, early adopters, a market, and — at last, awkwardly — measurement. Coaching ran this cycle in an unusually extreme form. By the time researchers began publishing serious tests of whether it works, coaching was already a global industry billing billions a year. It was installed in most large companies, with professional bodies, certification ladders, and a huge stock of testimonials standing in for an evidence base. The product was scaled first and tested later. So an entire generation of buying decisions was made on faith.
That era is over, and this article is about what replaced it. Between the early 2000s and the late 2010s, the does-it-work question finally went through the standard machinery of workplace science. First controlled studies. Then meta-analyses — studies that pool many trials into one estimate. Then systematic reviews of what strengthens or weakens the effects. The verdict is genuinely useful: more positive than sceptics expected, more conditional than the industry’s marketing implies.
What follows is the evidence in the order it arrived. First, the broad meta-analysis, then the workplace meta-analysis with its instructive moderators. Next, the honest longitudinal anchor and the mechanism studies. Then the critique of coaching’s favourite metric. Last, the uncomfortable scale lesson from the best-studied coaching literature of all — none of which was run on executives.
The boom that skipped the measurement
It is worth pausing on how odd the measurement gap was. Coaching’s claim is a causal one: structured one-to-one conversations change behaviour and performance. Causal claims about human development are exactly what organizational psychology knows how to test. The tools existed; what was missing for decades was any commercial pressure to use them. Coaching was sold relationship by relationship, and satisfaction was high. The buying executive was often also the person being coached — so nobody in the deal had an incentive to ask for a control group.
The result was a literature made of case studies, surveys of coachees who liked their coaches, and return-on-investment figures produced by the providers themselves. Practitioners could point to real theory underneath: goal-setting theory, behaviour-change models, the therapeutic-alliance research next door in psychotherapy. But the direct question — does workplace coaching cause measurable improvement against a comparison condition — sat unanswered while the industry grew around it.
When academic reviewers finally assembled the controlled studies, they were struck by how few existed relative to the size of the market (Theeboom, Beersma & van Vianen, 2014). The interesting part is what happened when that small literature was pooled. Coaching did not fail its exam. It passed — with grades worth reading closely.
The first real answer
The first broad meta-analysis of coaching in work settings came from Theeboom, Beersma and van Vianen in 2014. It pooled the controlled studies then available and sorted their outcomes into five domains: performance and skills, well-being, coping, work attitudes, and goal-directed self-regulation (Theeboom, Beersma & van Vianen, 2014).
The pooled effects were positive in every domain. Their sizes sat in the range the field calls moderate — roughly g ≈ 0.4 to 0.7, depending on the domain. The largest effects appeared for goal-directed self-regulation. Coachees became measurably better at setting specific goals, planning toward them, tracking their own progress, and adjusting. That is less glamorous than transformation, and more useful — self-regulation is the machinery any other change has to travel through. Performance and skill outcomes also showed solid effects; well-being and coping improved more modestly (Theeboom, Beersma & van Vianen, 2014).
The authors were properly careful about the small pool of studies. But the headline finding was clear and has held up: on average, coaching at work is effective, with effects large enough to matter and specific enough to point at a mechanism. A parallel meta-analysis published the following year reached a compatible conclusion. It added that the coach–coachee relationship itself predicted how strong the outcomes were — an early hint of where the causal machinery lives (Sonesh et al., 2015).
The workplace meta-analysis and its moderators
Two years later, Jones, Woods and Guillaume published the meta-analysis that most directly targets the corporate buying decision (Jones, Woods & Guillaume, 2016). It covered workplace coaching itself — employees coached in work settings — with learning and performance outcomes. It also had enough studies to test moderators: the conditions under which effects grow or shrink. The overall effect was again positive and roughly moderate. The moderator analyses are where the practical lessons sit. Three of them cut against expensive intuitions.
First, internal coaches did fine. In this meta-analysis, coaching delivered by internal coaches was at least as effective as coaching bought from external providers. That undercuts the idea that results scale with the day rate. Second, delivery format mattered less than expected. The analysis found no clear edge for face-to-face delivery over formats that blended in remote or tech-mediated elements — in 2016 a genuinely surprising result, and one that has aged rather well.
Third, and most surprising: coaching programs built around multisource (360-degree) feedback showed weaker effects than those without it (Jones, Woods & Guillaume, 2016). The authors’ reading was properly hedged. Multisource feedback may crowd the agenda with other people’s priorities rather than the coachee’s own goals. Or it may attach evaluative anxiety to what works best as a developmental relationship. Either way, the direction of the finding is a standing caution for the many programs that bolt coaching onto a 360 process by default. Across all three moderators the pattern points the same way: the active ingredients are in the process and the relationship, not the packaging.
The honest anchor
Meta-analyses pool many small studies. It steadies the picture to look closely at one large, careful, early study that the pooled numbers have to square with. The classic is Smither and colleagues’ longitudinal study — one that tracked people over time — of executive coaching inside a multinational (Smither, London, Flautt, Vargas & Kucine, 2003). Over a thousand senior managers were rated by bosses, peers and subordinates across two annual rounds of multisource feedback. Several hundred of them worked with an executive coach between the two.
The design was quasi-experimental rather than randomized — no coin flip decided who got a coach. But its scale, and its outcome measure — other people’s ratings, not the coachee’s own glow — make it the literature’s honest anchor. The result ran in coaching’s favour, and in modesty’s. Managers who worked with a coach improved more than their uncoached peers on multisource ratings. They were more likely to set specific rather than vague goals, and to ask their supervisors for ideas for improvement. And the improvement, while statistically reliable in a large sample, was small (Smither, London, Flautt, Vargas & Kucine, 2003).
That combination — real effect, small effect, big sample — is what a mature reading of coaching looks like. Coaching moves the needle on how a manager is actually experienced by the people around them. It does not transform anyone in a fiscal year. Programs that promise to transform people promise something the best long-term evidence has never shown. And programs dismissed for delivering “only” step-wise gains are being graded against a fantasy rather than the evidence.
What actually drives outcomes
If the average effect is moderate, the practical question becomes: what separates the engagements that work from those that do not? Here the literature borrows its central finding from psychotherapy research — common factors beat special ingredients. De Haan and colleagues studied executive-coaching outcomes across a large sample of coach–coachee pairs. The standout predictor was the strength of the working alliance: the coachee’s experience of the relationship as collaborative, goal-aligned and trusting. Coachee self-efficacy — belief in one’s own capacity to change — predicted outcomes too. The variables the market prices — the particular coach, the personality match, the coach’s preferred techniques — contributed far less (de Haan, Duckworth, Birch & Jones, 2013).
Bozer and Jones’ systematic review of what makes workplace coaching work organized the wider evidence into the same shape (Bozer & Jones, 2018). What recurs across studies: coachee motivation and self-efficacy, the quality of the coach–coachee relationship, and organizational support for the engagement. The coachee’s readiness does much of the work. Two absences from that list are as instructive as the entries. Coach certification and credentials — the industry’s main quality signal — are still largely unproven as predictors of outcome. And specific coaching methods, the branded frameworks that set providers apart, show little evidence of working better than one another.
The mechanism picture that emerges fits the meta-analyses. Coaching works when a motivated person, inside a relationship they trust, is held to a structured loop of specific goals, monitored behaviour and follow-through. Everything else is decoration, some of it expensive.
The ROI trap
Before the scale question, a detour through coaching’s favourite number — it still dominates sales decks. The industry’s traditional answer to the measurement problem was return on investment: headline multiples of the program’s cost. The multiple is built by asking coachees or sponsors to estimate the cash value of the changes, then dividing by the fee.
Grant’s critique of this practice is short and, once read, hard to unread: ROI as used in coaching is a poor and easily gamed metric (Grant, 2012). The cash-conversion step is subjective. The attribution step is heroic. The incentives of everyone producing the number point the same direction. And a metric that the person selling the service can inflate at will is not a metric.
His constructive point matters more than the demolition. Coaching’s defensible outcomes are the ones the controlled literature actually measures: goal attainment, well-being, engagement, observed behaviour change (Grant, 2012). Companies would judge coaching more honestly through that lens than by demanding a cash multiple no one can compute in good faith. The practical rule: treat any provider leading with an ROI multiple as making a marketing claim. Then move the conversation to pre-agreed behaviour and well-being outcomes, measured by someone other than the provider.
A coaching ROI multiple is usually a coachee’s monetized guess, attributed heroically and divided by the fee — by the party selling the engagement. A metric that can be inflated at will by its own vendor is not a metric (Grant, 2012). Evaluate against pre-specified goal-attainment and behaviour change instead.
The scale lesson from teacher coaching
The final piece of evidence comes from outside the corporate world, and it is the most rigorous coaching research there is. Teacher coaching — trained coaches working one-to-one with teachers on their classroom practice — has been tested in dozens of causal studies. Kraft, Blazar and Hogan pooled that causal evidence, and the effects dwarf most workplace programs: roughly 0.49 standard deviations on the quality of instruction, and roughly 0.18 on student achievement (Kraft, Blazar & Hogan, 2018). For a field intervention measured against hard outcomes, those are remarkable numbers.
≈0.49 SD The pooled causal effect of teacher coaching on instructional quality — with roughly 0.18 SD reaching student achievement — from the most rigorous coaching literature in existence (Kraft, Blazar & Hogan, 2018).
The finding that travels, though, is what happened to the numbers as programs grew. Effects were concentrated in smaller programs. When coaching models scaled up from dozens of teachers to hundreds or thousands, average effects shrank markedly (Kraft, Blazar & Hogan, 2018). The candidate explanations are mundane and universal. Scaled programs hire more coaches from a fixed supply of good ones. They dilute training and supervision, cut session frequency to stretch budgets, and enrol people with less appetite.
None of that is specific to schools. It is the economics of any labour-heavy, quality-sensitive service — which is precisely what corporate coaching is. Every company planning to spread coaching beyond the executive floor inherits this trade-off on day one. The thing that made the pilot work — scarce good coaches, generous dosage, volunteer coachees — is the thing that does not scale. The dosage–quality curve, not the pilot result, is what a scaled program should be designed against.
Does coaching work? A meta-analysis.Theeboom, Beersma & van Vianen, Journal of Positive Psychology, 2014
What the evidence doesn’t show
The coaching literature earned its positive verdict. But the verdict comes with limits, and buyers should hold them as firmly as the headline.
- Few active-control trials. Most controlled studies compare coaching to nothing — waitlists or no treatment — rather than to a cheaper active alternative such as structured goal-setting alone, training, or a good manager conversation. “Better than nothing” is established; “better than the alternatives per pound spent” mostly is not.
- Publication bias is plausible. The pooled literatures are small, effects vary widely across studies, and null coaching evaluations are exactly the kind of result that stays in a filing drawer — a caution the meta-analysts themselves raise (Theeboom, Beersma & van Vianen, 2014).
- Self-report inflates. Outcome domains lean on coachee-reported measures, which run warmer than observed behaviour; where other-rated outcomes anchor the estimate, effects are more modest (Smither, London, Flautt, Vargas & Kucine, 2003).
- Durability is unmeasured. Follow-ups typically end within months of the engagement. Whether coached gains persist at two or five years — or fade like most training effects without reinforcement — is essentially unknown.
- Executive coaching is the thin end of the evidence. The strongest pooled results come from workplace coaching broadly; the premium-priced executive tier rests on fewer and weaker studies than its market share implies (de Haan, Duckworth, Birch & Jones, 2013).
- Credentials are unvalidated. Certification, accreditation hours and branded methodologies — the industry’s quality signals — remain largely unstudied as predictors of outcomes (Bozer & Jones, 2018).
Where the evidence stops
- 1Few active-control trials
- 2Publication bias is plausible
- 3Self-report inflates
- 4Durability is unmeasured
- 5Executive coaching is the thin end of the evidence
- 6Credentials are unvalidated
Coaching by the evidence
Read as one body of work, the literature turns into a short operating manual — one that looks different from most corporate coaching programs now running.
Coach for specific goals with observable behaviour. The largest effects in the meta-analytic record sit on goal-directed self-regulation (Theeboom, Beersma & van Vianen, 2014). And the anchor study’s coached managers stood out precisely by setting specific rather than vague goals (Smither, London, Flautt, Vargas & Kucine, 2003). An engagement without a named behaviour to change is a pleasant conversation, not the intervention the evidence priced.
Select for coachee readiness, then protect the alliance. Motivation, self-efficacy and relationship quality are the recurring drivers (Bozer & Jones, 2018). The alliance predicts outcomes where coach pedigree does not (de Haan, Duckworth, Birch & Jones, 2013). Conscripted coachees and prestige coach-matching both spend money where the mechanism isn’t.
Do not default coaching onto 360 feedback, and keep the agenda owned by the coachee — the moderator evidence points that way, hedged but consistent (Jones, Woods & Guillaume, 2016). Use internal coaches without apology; the same evidence says they work.
Design for the dosage–quality curve before scaling. The strongest causal coaching literature shrank on contact with scale (Kraft, Blazar & Hogan, 2018). A rollout plan should fix session frequency, coach supervision and quality checks at full volume — not pilot volume. That is the difference between spreading coaching and diluting it.
Measure with pre-registered outcomes, not satisfaction or ROI. Decide before the program starts which behaviours and well-being measures count. Collect them from raters other than the coachee where possible, and refuse the cash-multiple theatre (Grant, 2012). Coaching survives honest measurement — that is the whole finding of this literature. It should be bought and run by people willing to apply some.
How Future Proof™ applies this.
The evidence says coaching works through goal-directed self-regulation, follow-through and alliance — and stops working when scale dilutes dosage. Future Proof’s AI coaching is built from exactly those mechanisms: every engagement runs on specific, named goals tied to observable behaviour; the Memory Coach schedules spaced follow-through checks so commitments resurface instead of evaporating; assessments and analytics measure the behaviour change itself, not satisfaction. Because the dosage is delivered by software, it does not thin as the program grows — the failure mode the scale evidence warns about — and human coaches are reserved for the judgment-heavy conversations where the alliance earns its cost, with the platform briefing them on what the data already shows.
See the platform →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above.
The evidence, by year
- 2003Smither
- 2012Grant
- 2013Haan
- 2014Theeboom
- 2015Sonesh
- 2016Jones
- 2018Bozer
- 2018Kraft
- Theeboom, T., Beersma, B., & van Vianen, A.E.M. (2014). Does coaching work? A meta-analysis on the effects of coaching on individual level outcomes in an organizational context. Journal of Positive Psychology 9(1): 1–18. PDF
- Jones, R.J., Woods, S.A., & Guillaume, Y.R.F. (2016). The effectiveness of workplace coaching: A meta-analysis of learning and performance outcomes from coaching. Journal of Occupational and Organizational Psychology 89(2): 249–277. DOI
- Sonesh, S.C., Coultas, C.W., Lacerenza, C.N., et al. (2015). The power of coaching: A meta-analytic investigation. Coaching: An International Journal of Theory, Research and Practice 8(2): 73–95. PDF
- Smither, J.W., London, M., Flautt, R., Vargas, Y., & Kucine, I. (2003). Can working with an executive coach improve multisource feedback ratings over time? Personnel Psychology 56(1): 23–44. PDF
- de Haan, E., Duckworth, A., Birch, D., & Jones, C. (2013). Executive coaching outcome research: The contribution of common factors such as relationship, personality match, and self-efficacy. Consulting Psychology Journal: Practice and Research 65(1): 40–57. PDF
- Grant, A.M. (2012). ROI is a poor measure of coaching success: Towards a more holistic approach using a well-being and engagement framework. Coaching: An International Journal of Theory, Research and Practice 5(2): 74–85. PDF
- Bozer, G., & Jones, R.J. (2018). Understanding the factors that determine workplace coaching effectiveness: A systematic literature review. European Journal of Work and Organizational Psychology 27(3): 342–361. PDF
- Kraft, M.A., Blazar, D., & Hogan, D. (2018). The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence. Review of Educational Research 88(4): 547–588. DOI
Coaching that survives scale.
Book a 20-minute demo. We’ll show you goal-directed coaching loops with spaced follow-through and measured behaviour change — at a dosage that doesn’t dilute when the program grows.