What actually predicts job performance.
Recruiters have argued about résumés, interviews, and tests for a hundred years. Personnel psychology kept score. A tour of the validity evidence — the 1998 league table, the 2022 re-ranking that dethroned the IQ test — and the principle behind Future Proof™ hiring campaigns.
The finding: After a century of validation research, the strongest predictors of job performance are methods that directly sample job-relevant knowledge and behaviour — structured interviews, job knowledge tests, work samples. The weakest are the proxies résumés are made of: years of education (~.10) and years of experience (~.18). A 2022 re-analysis lowered nearly every classic estimate and moved structured interviews to the top of the table.
The mechanism: Two forces explain the rankings. Behavioural consistency — the closer a measure sits to the actual work, the better it forecasts the work. And a statistical one: the famous 1998 estimates were inflated by overly aggressive corrections for range restriction; redoing the corrections demoted general cognitive ability from .51 to roughly .31.
The product: Future Proof hiring campaigns lead with structured, job-relevant assessment — subject batteries and work-relevant questions, scored identically for every candidate — instead of screening on résumé proxies.
In this article
- 01Eighty-five years, one league table
- 02The 2022 re-analysis: same data, different corrections
- 03What a .42 is actually worth
- 04What the strong predictors share
- 05Why the winners win: behavioural consistency
- 06Where résumé proxies land
- 07What the evidence doesn’t show
Hiring debates run hot because everyone has data — a career of hires remembered selectively, with the misses explained away and the great picks credited to judgment. What almost no single career supplies is a scored forecast: predictions recorded before outcomes, outcomes measured independently, across enough cases to separate skill from noise. One field has been running exactly that ledger since before the Second World War. Its totals are the closest thing hiring has to settled fact.
Every hiring method is a prediction instrument. A résumé screen, a panel interview, a coding exercise, a reference call — each is a bet that some signal you can see before the offer letter forecasts behaviour after it. Which turns hiring into an unusually answerable scientific question: which instruments predict best?
Personnel psychology has been keeping score for roughly a century. Its unit of account is the validity coefficient — the link between scores on a hiring method and later job performance, usually a manager’s rating. One study alone is small and noisy. So the field leans on meta-analysis: pool hundreds of studies, and estimate what each method is worth on average.
Two papers bracket the modern conversation. The first, Schmidt and Hunter’s synthesis of 85 years of findings (Schmidt & Hunter, 1998), became one of the most cited papers in applied psychology and hardened into the field’s league table. The second, a re-analysis by Sackett, Zhang, Berry, and Lievens (Sackett et al., 2022), reopened the case a quarter-century later — and changed the rankings.
Eighty-five years, one league table
The first serious attempt to rank predictors at scale was Hunter and Hunter’s 1984 synthesis (Hunter & Hunter, 1984). It concluded that general mental ability (GMA) — in practice, scores on cognitive ability tests — was the most valid single predictor available across the full range of jobs. Schmidt and Hunter’s 1998 paper pulled the evidence together across 19 selection procedures (Schmidt & Hunter, 1998). Its headline estimates became canon:
- Work sample tests: .54, the top single predictor in the table.
- GMA tests: .51 against overall job performance.
- Structured interviews: .51 — versus .38 for unstructured interviews.
- Job knowledge tests: .48.
- Conscientiousness measures: .31; reference checks .26.
- Years of job experience: .18; years of education: .10.
- Graphology: .02, and age: −.01 — essentially chance.
The table’s bottom rows deserve as much attention as its top, because they price rituals still in daily use. Graphology at .02 remains the standing reminder that an industry can sell a zero-validity instrument for decades. Age at −.01 prices the quiet assumption behind every “too junior/too senior” screen. A league table’s job is exactly this — putting beloved practices and effective ones on one axis, where the comparisons cannot be politely avoided.
Two further conclusions mattered as much as the rankings. First, combinations: pairing a GMA test with a structured interview or a work sample produced the highest composite validities in the analysis (Schmidt & Hunter, 1998). Second, the spread from top to bottom of the table is enormous — the difference between an instrument with a real predictive edge and an expensive ritual. For two decades, this table was the standard answer to “how should we hire?”
The 2022 re-analysis: same data, different corrections
Meta-analytic validities are not raw correlations. They are corrected — for measurement error in the performance yardstick, and for range restriction. Range restriction means validation samples contain only people who were already hired, so the observed correlation understates what the predictor would do across a full applicant pool. The corrections are legitimate in principle. The question is how large they should be.
What happened next is science working as advertised, and rarer in applied fields than anyone admits. The discipline’s most famous numbers were challenged not by outsiders with a rival product, but by senior insiders re-auditing the arithmetic. Sackett and colleagues argued that the corrections behind the classic estimates were systematically too aggressive (Sackett et al., 2022). Most validation studies are concurrent — run on current employees rather than applicant groups tracked over time. In that design, the data needed to estimate applicant-pool restriction is usually missing. The classic analyses, they argued, applied corrections the underlying study designs could not justify — inflating some methods far more than others.
Re-estimated with corrections the authors considered defensible, the table compressed and reshuffled. Structured interviews moved to the top at roughly .42, followed by job knowledge tests (about .40), empirically keyed biodata (about .38), and work samples (about .33). GMA fell from .51 to roughly .31, and conscientiousness from .31 to about .19 (Sackett et al., 2022).
Read this carefully: nothing here makes cognitive ability useless. A correlation near .31 is still large by selection standards, and still cheap to get. What changed is the relative picture. In the 1998 table, general ability sat above every job-specific method. In the 2022 table, methods that directly sample job-relevant knowledge and behaviour sit above it.
What a .42 is actually worth
Validity coefficients look small to anyone calibrated on physics, so their economics deserve translation. A correlation of .42 between a selection composite and job performance means this: move from hiring at random to hiring the top scorers, and the expected performance of each hire shifts by a substantial fraction of a standard deviation. And the standard deviation of employee output, in the utility literature both papers inherit, is worth a large slice of salary each year (Schmidt & Hunter, 1998). Multiply a per-hire gain across a year’s hiring volume and a tenure of years. The difference between a .42 process and a .19 process is not a psychometric nicety. It is one of the larger unmanaged line items in most companies’ economics.
The translation also disciplines hope in the other direction. Even the best composite leaves most of the variance in performance unexplained. Jobs are noisy, situations dominate, and people change. So no process, however evidence-based, picks winners reliably one person at a time. The honest promise of a high-validity process is statistical: better hires on average, fewer disasters on average, compounding across volume. Any vendor promising individual certainty has left the literature; any executive who dismisses a .42 because one great hire once interviewed badly has misread what the number claims.
The field’s most famous numbers rested on corrections the underlying studies could not support.The central claim of Sackett et al. 2022, paraphrased
What the strong predictors share
The reshuffled table has a legible pattern. The methods at the top — structured interviews, job knowledge tests, work samples, biodata keyed against actual outcomes — all sample job-relevant behaviour or knowledge directly, rather than inferring it from credentials. This is the behavioural-consistency principle: the best predictor of performance on a task is performance on a similar task, observed under standardized conditions. Even the classic work-sample estimate of .54 was revised downward by later, more careful meta-analysis (Roth et al., 2005). Yet the sample-the-work family remains at or near the top of every ranking produced since.
Beneath the league table run three patterns that survive every correction regime. They are worth extracting, because they generalize to instruments the table never rated. The interview literature is the cleanest illustration of the second: structure beats intuition. Meta-analyses in the 1990s established that interviews built from job analysis — the same questions asked of every candidate, answers scored against anchored criteria — are substantially more valid than free-form conversations (McDaniel et al., 1994). Validity also rises with the degree of structure (Huffcutt & Arthur, 1994).
The unstructured interview is not merely weaker; it can actively mislead. In one set of experiments, interviewers predicted students’ upcoming semester GPA. Adding an unstructured interview to background facts made the predictions worse than the facts alone. And when some interviewees secretly answered questions at random, interviewers rarely noticed anything amiss (Dana et al., 2013).
The unstructured interview can be worse than nothing: adding one to background information made predictions less accurate than the background information alone — the conversation manufactures confidence without adding signal (Dana et al., 2013).
A third pattern: how signals are combined matters as much as which signals are collected. One meta-analysis compared holistic expert judgment against mechanical, rule-based combination of the same information. The rule-based approach was more accurate, consistently, in selection and admissions decisions (Kuncel et al., 2013). Gut-feel blending of good data gives back much of the validity the data provided.
Collect structured signals, then combine them by rule. Mechanical combination of the same information consistently out-predicted holistic expert judgment in selection and admissions — the final gut-feel synthesis is where good data goes to lose its validity (Kuncel et al., 2013).
Personality measures sit in a middle tier. Conscientiousness is the one Big Five trait that predicts performance across nearly all job families (Barrick & Mount, 1991). But its validity is modest, and the 2022 re-analysis lowered it further (Sackett et al., 2022). Personality is a useful add-on to job-relevant measurement, not a substitute for it.
Why the winners win: behavioural consistency
The re-ranked table has a theme that no single coefficient states. The methods now at the top — structured interviews probing past job behaviour, knowledge tests sampling what the work requires knowing, work samples staging the work itself — all run on one principle: the best predictor of behaviour is behaviour of the same kind. Each shortens the inferential distance between what is measured and what is forecast. General ability sits lower not because thinking doesn’t matter. Its route to job performance runs through middle steps — chiefly the acquisition of job knowledge — that the top methods measure directly (Hunter & Hunter, 1984). The 2022 corrections did not overturn behavioural consistency; they revealed it more clearly, by deflating the one general-purpose instrument the old corrections had flattered most (Sackett et al., 2022).
The principle is also the practical design compass, because it answers questions no league table covers. When a role is novel and no validity study exists, build the instrument that samples the actual work — the case, the scenario, the knowledge the first ninety days will demand. Built that way, it will inherit the top of the table’s logic even before local data accumulates. Proxies, by the same compass, are always a bet that distance doesn’t matter; the century of scorekeeping is the record of that bet losing.
Where résumé proxies land
Which brings the tour to the instrument that actually runs most of the world’s hiring — not any test or interview, but the document that decides who ever reaches one. The awkward finding is stable across both the 1998 and 2022 analyses: the signals that dominate real-world screening are among the weakest in the table. Years of education correlates with performance at roughly .10, years of job experience at roughly .18 (Schmidt & Hunter, 1998). And the predictive value of experience concentrates in the early years on a job, rather than growing without limit (Schmidt & Hunter, 1998). A résumé is, for the most part, a list of exactly these proxies.
The literature’s verdict is not that résumés are worthless. It is that a funnel which filters hard on proxies before ever measuring job-relevant knowledge or behaviour discards most of the predictive power available to it. The arithmetic is brutal at scale. A screen that cuts ninety percent of applicants on .10-validity signals has made its biggest decision — who gets measured at all — with its least valid instrument, and no downstream rigor can recover the strong candidates it silently removed. Reordering the funnel, so that cheap structured measurement comes before the proxy filter rather than after it, is the single largest validity upgrade most hiring systems have available. The modern assessment economics our skills-based-hiring review describes is what finally made the reordering affordable.
.10 / .18 The validity of years-of-education and years-of-experience — the two signals most résumé screens filter hardest on, sitting near the bottom of the table under every correction regime (Schmidt & Hunter, 1998).
How Future Proof™ applies this.
Hiring campaigns on the platform lead with the top of the validity table instead of the bottom. Candidates start with subject batteries — job knowledge tests scoped to the specific role and stack — and work-relevant questions built from the role’s actual competencies, delivered and scored identically for every candidate. Résumé fields don’t gate who gets measured: the structured assessment comes first, and comparable, job-relevant scores come out the other side.
See hiring campaigns →What the evidence doesn’t show
This literature is unusually deep. Because people cite it with so much confidence, its limits are worth stating plainly:
- The corrections debate is not settled. The 2022 downward revision has itself been challenged — Oh, Le, and Roth argue that abandoning range-restriction corrections in concurrent studies overcorrects in the opposite direction (Oh et al., 2023). The exact magnitudes remain contested; the shift in rank order has proven more robust than any single number.
- These are averages, not guarantees. A meta-analytic .42 for structured interviews describes well-built, consistently scored interviews. It says nothing about a loosely scripted panel that shares the label. Validity varies with job complexity and with the quality of execution.
- The criterion is soft. Most validation studies use supervisor ratings as the performance measure. Ratings are themselves unreliable and partially contaminated, so “predicts job performance” mostly means “predicts how bosses rate performance” — related to, but not identical with, objective output.
- Validity is not a complete hiring policy. Predictor choice interacts with subgroup score differences, adverse impact, cost, and candidate experience. A validity ranking is one input to selection-system design, not the whole design.
- Generalization has edges. The evidence base is dominated by Western labour markets, incumbent samples, and typical rather than maximal performance. Effects in other contexts are plausible but less directly established.
Where the evidence stops
- 1The corrections debate is not settled
- 2These are averages, not guarantees
- 3The criterion is soft
- 4Validity is not a complete hiring policy
- 5Generalization has edges
None of these caveats rescues the bottom of the table. Under every correction regime yet proposed, structured, job-relevant measurement out-predicts credential proxies. The century-long argument is now about how large the gap is — not about which direction it points, and not about whether the funnel most companies run points the other way.
Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
The evidence, by year
- 1984Hunter
- 1991Barrick
- 1994McDaniel
- 1994Huffcutt
- 1998Schmidt
- 2005Roth
- 2013Dana
- 2013Kuncel
- 2022Sackett
- 2023Oh
- Schmidt, F.L., & Hunter, J.E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin 124(2): 262–274. DOI
- Hunter, J.E., & Hunter, R.F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin 96(1): 72–98. DOI
- McDaniel, M.A., Whetzel, D.L., Schmidt, F.L., & Maurer, S.D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology 79(4): 599–616. DOI
- Huffcutt, A.I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology 79(2): 184–190. DOI
- Barrick, M.R., & Mount, M.K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology 44(1): 1–26. DOI
- Roth, P.L., Bobko, P., & McFarland, L.A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology 58(4): 1009–1037. PDF
- Dana, J., Dawes, R., & Peterson, N. (2013). Belief in the unstructured interview: The persistence of an illusion. Judgment and Decision Making 8(5): 512–520. PDF
- Kuncel, N.R., Klieger, D.M., Connelly, B.S., & Ones, D.S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology 98(6): 1060–1072. PDF
- Oh, I.-S., Le, H., & Roth, P.L. (2023). Revisiting Sackett et al.’s (2022) rationale behind their recommendation against correcting for range restriction in concurrent validation studies. Journal of Applied Psychology 108. PDF
Run a hiring funnel that measures what predicts.
Book a 20-minute demo and we’ll build a campaign for one of your real roles — subject battery, work-relevant questions, structured scoring — and show you what the funnel looks like when assessment comes before the résumé sort.