When formulas beat experts.
Gather good assessment information, then hand the final combination to expert intuition — and validity measurably drops. Seventy years of evidence on mechanical versus clinical prediction, the psychology of why it is still ignored, and how Future Proof™ structures the last mile of a hiring decision.
The finding: When the same information is available, combining it by explicit rule (weights, sums, formulas) predicts outcomes at least as well as expert judgment in the overwhelming majority of studies. In selection and admissions specifically, mechanical combination beats holistic expert combination by roughly 50% more predictive validity. The result has replicated across seven decades, hundreds of studies, and nearly every domain tested.
The mechanism: Human combiners are inconsistent — the same profile gets different verdicts on different days — and they overweight vivid, recent, and interview-borne impressions. A formula applies the same policy every time; consistency alone explains most of its edge. Even models of a judge’s own policy beat the judge.
The product: Future Proof keeps humans in charge of what to measure and what to value — and gives the arithmetic to the machine: structured scorecards, preregistered weights, mechanical composites, and documented overrides that are tracked and audited.
In this article
- 01Meehl’s little book and its long shadow
- 02Why the formula wins: consistency is a superpower
- 03What the lens-model tradition adds
- 04The professional resistance
- 05Running the rule inside a real organization
- 06What this does and doesn’t say about AI hiring
- 07What the evidence doesn’t show
- 08What this means for practice
Of all the findings in this library, this one has the strangest status. It is among the most replicated results in applied psychology — and among the least used. Companies that would never field an unvalidated test happily run the last, highest-stakes step of selection on a method the evidence retired in 1954. Understanding why — and what the alternative costs, which is nearly nothing — is the purpose of what follows.
Every selection process ends the same way. A table of evidence — scores, interview ratings, work samples, references — and a person or committee looking at it, weighing it, and pronouncing. That final act of weighing feels like the moment expertise matters most. The hiring manager synthesizes; the committee deliberates; professional judgment assembles the “whole picture”.
It is also the moment, the evidence says, where the process gives back much of what the earlier rigor earned. A firm can spend heavily on validated assessments and trained interviewers. Then, in the final meeting, it recombines those signals by feel — quietly trading measurable validity for the felt experience of judgment. The waste is invisible because no single decision reveals it. It shows up only in the statistics, in outcomes across many hires that nobody links back to the combining step.
The uncomfortable distinction this article covers is between collecting information and combining it. Expert judgment is often excellent at the first — knowing what to look for, drawing it out, scoring a structured interview. The evidence about the second is among the most consistent in applied psychology. When the information is on the table, combining it by explicit rule beats combining it by intuition. And the profession that most needs to know this has spent seventy years declining to act on it.
Meehl’s little book and its long shadow
The controversy has a birth certificate: a slim 1954 book by Paul Meehl. He assembled every study comparing “clinical” prediction — an expert weighing case information in their head — against “statistical” prediction, the same information combined by formula. The comparisons covered academic success, parole outcomes, and psychiatric prognosis. They kept coming out the same way: the formula matched or beat the expert (Meehl, 1954). Meehl, himself a practicing clinician, spent the rest of his career noting that the finding was as replicated as anything in psychology — and as ignored.
The reception was everything the finding was not: heated, personal, and unresolved for decades. Clinicians read the book as an attack on their craft. Meehl kept insisting it was an attack on one step of their craft — the final arithmetic. On his account, everything else clinicians do remained untouched and vital: drawing out information, building rapport, choosing what to measure. The distinction never quite landed. The same category error still powers the debate’s modern revivals: defenders of judgment defend the whole profession against a finding that only ever concerned the adding-up.
The evidence since has only widened the ledger. The most complete meta-analysis of prediction comparisons pooled 136 studies across medicine, mental health, education, and personnel. It found mechanical prediction equal or better in the vast majority (Grove, Zald, Lebow, Snitz & Nelson, 2000). Clinical judgment won rarely and by small margins — mostly when the clinician had information the formula lacked.
And the domain that concerns this library got its own dedicated synthesis. Across selection and admissions studies, mechanical combination of assessment data outpredicted holistic expert combination of the same data by a wide margin (Kuncel, Klieger, Connelly & Ones, 2013). The gain from switching came to roughly half again as much validity from identical information.
≈ +50% The validity gained in selection and admissions by switching the final step from holistic to mechanical combination — same instruments, same information, different arithmetic (Kuncel et al., 2013).
Why the formula wins: consistency is a superpower
Before asking why the formula wins, clear away one suspicion: that it wins through sophistication — that somewhere in the studies a powerful statistical model is out-thinking the humans. The opposite is true, and that fact is the finding’s teeth. Nothing about the result requires formulas to be clever.
The decisive experiments are the ones using improper models — unit weights, or even models built to mimic a specific judge’s own policy. Give every predictor equal weight and the crude sum still beats the expert who insists each case needs bespoke weighting (Dawes, 1979). Stranger still, a regression model of a single judge — built solely from that judge’s past ratings — outpredicts the judge it was modeled on. The result is known as bootstrapping. The judge’s own policy, applied with perfect consistency, is better than the judge.
The victory belongs to consistency, not sophistication: unit-weighted sums beat bespoke expert weighting, and a model of a judge outpredicts the judge it was built from (Dawes, 1979). Nothing about the fix requires clever statistics — only the same policy, every time.
That locates the human deficit precisely: not in the values or the insight, but in the application. Human combiners are noisy. The same file read on Tuesday and Friday earns different verdicts; order effects, fatigue, mood, and the vividness of the last interview all leak into the weighting.
Humans are also seduced by certain channels, and the seduction is systematic. Broad reviews of selection judgment show a persistent pattern: judges overweight unstructured interview impressions — the least valid common signal — at the expense of the most valid ones (Highhouse, 2008). The reason: the impression arrives wrapped in narrative and eye contact. A formula has no Friday, no narrative, and no favorite candidate. It simply applies the agreed policy, every time — and in a noisy world, that is worth more than brilliance.
What the lens-model tradition adds
A second research tradition arrived at the same place by a different road. Its findings dissolve the last defense of intuitive combination — the belief that expert synthesis is too rich to model. This program spent decades modeling expert judges directly, capturing from a judge’s actual decisions the implicit weights they place on each cue. Its meta-analytic summary is a portrait of professional judgment as it actually works: simple linear mixes of a handful of cues capture most expert judgment policies well (Karelaia & Hogarth, 2008).
Judges believe they use complex, configural reasoning that their own data do not show. And the ceiling on judgment accuracy is set mostly by the environment’s predictability and the judge’s consistency — not by insight. The romantic account of expert synthesis — dozens of subtle factors, weighed in ineffable combination — is not what the measurements find. They find three to five cues, weighted noisily.
This is liberating rather than deflating, because it means the expert’s real contribution can be captured and kept. Ask your best hiring managers what they look for. Extract the cues and weights their judgments actually reflect, and encode that policy. At that point every candidate gets the best version of that manager’s judgment — including the ones judged at 6 p.m. on a Friday. The expertise is honored; the noise is retired.
The professional resistance
The finding’s persistence makes its neglect the more interesting fact. If the evidence is this old and this one-sided, the refusal to adopt it needs explaining too. The field has studied that as well — turning the same research tools on the deciders that it once turned on the decisions. Practitioners raise the same objections across decades: my cases are unique; formulas miss the intangibles; judgment is my professional craft. The objections have been examined and found wanting (Dawes, Faust & Meehl, 1989).
The uniqueness claim dissolves under a simple observation: formulas beat experts case by case, not just on average. The intangibles, once specified, can simply be scored and added to the model. What remains is the psychology. Combining by rule feels like giving up expertise, and organizations defer to that feeling because the cost — validity forgone — is statistical, delayed, and invisible in any single hire.
The modern version of the resistance has a name: algorithm aversion. In controlled studies, people who watch an algorithm err abandon it at far higher rates than they abandon equally erring humans — even when the algorithm performs better overall (Dietvorst, Simmons & Massey, 2015). The machine’s mistakes are treated as disqualifying in a way human mistakes never are. The same work found the practical antidote: people tolerate and use algorithms far more willingly when given even a modest ability to adjust the output. Governance design, not persuasion, is what closes the adoption gap.
Algorithm aversion is asymmetric: people who watch a formula err abandon it at far higher rates than they abandon an equally erring human (Dietvorst, Simmons & Massey, 2015). Build the tolerance in — a modest, documented ability to adjust the output keeps deciders using the rule.
Running the rule inside a real organization
The rollout problems are political before they are technical, and the successful deployments share a shape. The weights are set collectively and in advance — by the hiring team, against the job analysis, before any candidate is seen. That converts the formula from an imposition on managers into a record of their own considered policy. The composite is presented as the default with a documented exit, not a mandate: any decider may override, in writing, with the reason logged. And the loop is closed annually. Composite-followed and composite-overridden decisions are compared against outcomes, in aggregate, without naming names.
Each element earns its place. Advance agreement removes the suspicion that the weights were tuned toward a favored candidate. The documented exit respects the genuine broken-leg case while taxing the hollow one — writing “I have a feeling” under a timestamp is its own deterrent. And the annual audit turns the philosophical argument into a local, empirical one. After a year, the organization is no longer debating Meehl. It is reading its own override hit rate — which, in most reported experience, does the persuading that seventy years of literature could not (Dawes, Faust & Meehl, 1989).
Committees deserve a special note, because the group version of clinical combination adds problems of its own. The loudest voice becomes the de facto weights; early opinions anchor later ones; the felt legitimacy of consensus stands in for validity. The mechanical alternative does not abolish the meeting; it repurposes it. The meeting’s job becomes auditing the inputs — was this interview scored properly, is this test result odd — rather than re-deriving the arithmetic by charisma.
What this does and doesn’t say about AI hiring
No discussion of formulas deciding about people can end in 1954. The word “algorithm” now carries freight Meehl never imagined, and both the boosters and the critics routinely misread this literature as settling questions it never asked. It gets dragged into debates about algorithmic hiring, so its boundaries need drawing.
The mechanical-combination evidence concerns transparent rules over validated, job-relevant inputs — scores from structured assessments whose own validity is established. It does not bless opaque models trained on historical hiring outcomes. Those models inherit whatever the history contains: proxies, imbalances, and the recorded preferences of past deciders. The auditing literature on modern hiring algorithms documents exactly those risks and the diligence they demand (Raghavan, Barocas, Kleinberg & Levy, 2020).
The defensible position keeps the two questions separate. What goes into the decision is a validity-and-fairness question, settled by evidence per input. How the inputs are combined is the question this article settles — by rule, transparently, with the weights written down where they can be examined and contested. Opacity is not mechanical combination’s ally; it is its corruption.
There is no controversy in social science which shows such a large body of qualitatively diverse studies coming out so uniformly in the same direction.Paul Meehl, reflecting on four decades of clinical-versus-statistical comparisons.
What the evidence doesn’t show
- It doesn’t remove humans from selection. Everything upstream of combination — job analysis, choosing predictors, conducting structured interviews, scoring work samples — is human work the formula depends on. The finding concerns one step: the arithmetic at the end (Kuncel et al., 2013).
- It doesn’t forbid overrides. The “broken leg” case — decisive information the model never sees — is real and rare. The evidence-based policy allows overrides, requires them to be documented, and audits their hit rate, which in practice teaches organizations how rare genuine broken legs are (Dawes et al., 1989).
- It doesn’t require sophisticated models. Unit-weighted sums capture most of the gain; the victory belongs to consistency, not to statistical complexity (Dawes, 1979).
- It doesn’t validate any particular vendor’s black box. A model no one can inspect fails the transparency condition that makes mechanical combination auditable and contestable (Raghavan et al., 2020).
Where the evidence stops
- 1It doesn’t remove humans from selection
- 2It doesn’t forbid overrides
- 3It doesn’t require sophisticated models
- 4It doesn’t validate any particular vendor’s black box
What this means for practice
The reform this literature asks for is unusually cheap, which is part of what makes its neglect expensive. No new instruments, no new vendors, no extra candidate time — just a different final step performed on evidence already collected.
Split your selection process at the seam the evidence identifies. Upstream, invest in the human craft: define what the role requires, choose validated instruments, train interviewers, build anchored scorecards. At the seam, write the policy down before seeing candidates — which signals count, with what weights, combined how. A one-page decision rule, agreed while heads are cool. Downstream, let the rule run: compute the composite, rank within error bands, and treat the output as the decision’s default.
Then govern the exceptions instead of pretending they won’t occur. Allow overrides; require a written reason at the moment of override; review the overrides annually against outcomes. Organizations that run this loop discover two things. Override hit rates are humbling, and the write-it-down requirement alone extinguishes most overrides. The final mile of hiring is where organizations pay for the theater of synthesis with the substance of validity. The fix costs a page of arithmetic and a little professional humility — and it has been sitting in the literature, fully replicated, since before most current executives were born.
How Future Proof™ applies this: humans set the policy, the engine runs it.
Hiring campaigns in the platform end in a mechanical composite, not a vibe: every signal — assessments, structured interview ratings, scenario scores — enters a scorecard whose weights were configured before the first candidate applied. The arithmetic is transparent and inspectable; rankings respect measurement error rather than decimal theater; and overrides are supported the way the literature prescribes — documented at the moment, tracked over time, reported against outcomes. Expertise decides what matters. The engine makes sure it matters the same way every time.
See structured scorecards →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
The evidence, by year
- 1954Meehl
- 1979Dawes
- 1989Dawes
- 2000Grove
- 2008Highhouse
- 2008Karelaia
- 2013Kuncel
- 2015Dietvorst
- 2020Raghavan
- Meehl, P.E. (1954). Clinical versus Statistical Prediction: A Theoretical Analysis and a Review of the Evidence. University of Minnesota Press. PDF
- Grove, W.M., Zald, D.H., Lebow, B.S., Snitz, B.E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment 12(1): 19–30. PDF
- Kuncel, N.R., Klieger, D.M., Connelly, B.S., & Ones, D.S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology 98(6): 1060–1072. PDF
- Dawes, R.M. (1979). The robust beauty of improper linear models in decision making. American Psychologist 34(7): 571–582. PDF
- Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology 1(3): 333–342. PDF
- Dawes, R.M., Faust, D., & Meehl, P.E. (1989). Clinical versus actuarial judgment. Science 243(4899): 1668–1674. PDF
- Dietvorst, B.J., Simmons, J.P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General 144(1): 114–126. PDF
- Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating bias in algorithmic hiring: Evaluating claims and practices. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency: 469–481. PDF
- Karelaia, N., & Hogarth, R.M. (2008). Determinants of linear judgment: A meta-analysis of lens model studies. Psychological Bulletin 134(3): 404–426. PDF
Write the weights down before the candidates arrive.
Book a 20-minute demo. We’ll show you preregistered scorecards, transparent composites, and override tracking — the last mile of hiring, run the way the evidence says.