Research · Assessment Science
Assessment Science · Personnel Selection

What actually predicts job performance.

Recruiters have argued about résumés, interviews, and tests for a hundred years. Personnel psychology kept score. A tour of the validity evidence — the 1998 league table, the 2022 re-ranking that dethroned the IQ test — and the principle behind Future Proof™ hiring campaigns.

TL;DR

The finding: After a century of validation research, the strongest predictors of job performance are methods that directly sample job-relevant knowledge and behaviour — structured interviews, job knowledge tests, work samples. The weakest are the proxies résumés are made of: years of education (~.10) and years of experience (~.18). A 2022 re-analysis lowered nearly every classic estimate and moved structured interviews to the top of the table.

The mechanism: Two forces explain the rankings. Behavioural consistency — the closer a measure sits to the actual work, the better it forecasts the work. And a statistical one: the famous 1998 estimates were inflated by overly aggressive corrections for range restriction; redoing the corrections demoted general cognitive ability from .51 to roughly .31.

The product: Future Proof hiring campaigns lead with structured, job-relevant assessment — subject batteries and work-relevant questions, scored identically for every candidate — instead of screening on résumé proxies.

Every hiring method is a prediction instrument. A résumé screen, a panel interview, a coding exercise, a reference call — each is a bet that some signal observable before the offer letter forecasts behaviour after it. Which turns hiring into an unusually answerable scientific question: which instruments predict best?

Personnel psychology has been keeping score for roughly a century. Its unit of account is the validity coefficient — the correlation between scores on a selection method and later job performance, usually measured by supervisor ratings. Because individual validation studies are small and noisy, the field leans on meta-analysis: pooling hundreds of studies to estimate what each method is worth on average.

Two papers bracket the modern conversation. The first, Schmidt and Hunter’s synthesis of 85 years of findings (Schmidt & Hunter, 1998), became one of the most cited papers in applied psychology and hardened into the field’s league table. The second, a re-analysis by Sackett, Zhang, Berry, and Lievens (Sackett et al., 2022), reopened the case a quarter-century later — and changed the rankings.

Eighty-five years, one league table

The first serious attempt to rank predictors at scale was Hunter and Hunter’s 1984 synthesis, which concluded that general mental ability (GMA) — in practice, scores on cognitive ability tests — was the most valid single predictor available across the full range of jobs (Hunter & Hunter, 1984). Schmidt and Hunter’s 1998 paper consolidated the evidence across 19 selection procedures (Schmidt & Hunter, 1998), and its headline estimates became canon:

  • Work sample tests: .54, the top single predictor in the table.
  • GMA tests: .51 against overall job performance.
  • Structured interviews: .51 — versus .38 for unstructured interviews.
  • Job knowledge tests: .48.
  • Conscientiousness measures: .31; reference checks .26.
  • Years of job experience: .18; years of education: .10.
  • Graphology: .02, and age: −.01 — essentially chance.

Two further conclusions mattered as much as the rankings. First, combinations: pairing a GMA test with a structured interview or a work sample produced the highest composite validities in the analysis (Schmidt & Hunter, 1998). Second, the spread from top to bottom of the table is enormous — the difference between an instrument with a real predictive edge and an expensive ritual. For two decades, this table was the standard answer to “how should we hire?”

The 2022 re-analysis: same data, different corrections

Meta-analytic validities are not raw correlations. They are corrected — for measurement error in the performance criterion, and for range restriction: validation samples contain only people who were already hired, so the observed correlation understates what the predictor would do across a full applicant pool. The corrections are legitimate in principle. The question is how large they should be.

Sackett and colleagues argued that the corrections behind the classic estimates were systematically too aggressive (Sackett et al., 2022). Most validation studies are concurrent — run on incumbent employees rather than tracked applicant cohorts — and in that design the information needed to estimate applicant-pool restriction is usually absent. The classic analyses, they argued, applied corrections the underlying study designs could not justify, inflating some methods far more than others.

Re-estimated with corrections the authors considered defensible, the table compressed and reshuffled. Structured interviews moved to the top at roughly .42, followed by job knowledge tests (about .40), empirically keyed biodata (about .38), and work samples (about .33). GMA fell from .51 to roughly .31, and conscientiousness from .31 to about .19 (Sackett et al., 2022).

Read this carefully: nothing here makes cognitive ability useless. A correlation near .31 is still large by selection standards, and still cheap to obtain. What changed is the relative picture. In the 1998 table, general ability sat above every job-specific method; in the 2022 table, methods that directly sample job-relevant knowledge and behaviour sit above general ability.

Schmidt & Hunter 1998 Sackett et al. 2022 0 .10 .20 .30 .40 .50 .60 Structured interview .51 .42 Job knowledge test .48 .40 Biodata (keyed) .35 .38 Work sample .54 .33 Cognitive ability (GMA) .51 .31 Conscientiousness .31 .19
Figure 1. Estimated operational validity against overall job performance: classic estimates (coral) vs the 2022 re-estimates (indigo). Values from Schmidt & Hunter 1998 and Sackett et al. 2022. Note both the lower magnitudes and the new rank order.
The field’s most famous numbers rested on corrections the underlying studies could not support. The central claim of Sackett et al. 2022, paraphrased

What the strong predictors share

The reshuffled table has a legible pattern. The methods at the top — structured interviews, job knowledge tests, work samples, biodata keyed against actual outcomes — all sample job-relevant behaviour or knowledge directly, rather than inferring it from credentials. This is the behavioural-consistency principle: the best predictor of performance on a task is performance on a similar task, observed under standardized conditions. Even the classic work-sample estimate of .54 was revised downward by later, more careful meta-analysis (Roth et al., 2005), yet the sample-the-work family remains at or near the top of every ranking produced since.

The interview literature is the cleanest illustration of a second pattern: structure beats intuition. Meta-analyses in the 1990s established that interviews built from job analysis — the same questions asked of every candidate, answers scored against anchored criteria — are substantially more valid than free-form conversations (McDaniel et al., 1994), and that validity rises with the degree of structure (Huffcutt & Arthur, 1994). The unstructured interview is not merely weaker; it can actively mislead. In experiments where interviewers predicted students’ upcoming semester GPA, adding an unstructured interview to background information made predictions worse than the background information alone — and when some interviewees covertly answered questions at random, interviewers rarely noticed anything amiss (Dana et al., 2013).

A third pattern: how signals are combined matters as much as which signals are collected. A meta-analysis comparing holistic expert judgment against mechanical, rule-based combination of the same information found the mechanical approach consistently more accurate in selection and admissions decisions (Kuncel et al., 2013). Gut-feel synthesis of good data gives back much of the validity the data provided.

Personality measures occupy a middle tier. Conscientiousness is the one Big Five trait that predicts performance across essentially all occupational groups (Barrick & Mount, 1991), but its validity is modest, and the 2022 re-analysis lowered it further (Sackett et al., 2022). Personality is a useful supplement to job-relevant measurement, not a substitute for it.

Where résumé proxies land

The awkward finding — stable across both the 1998 and 2022 analyses — is that the signals dominating real-world screening are among the weakest in the table. Years of education correlates with performance at roughly .10, years of job experience at roughly .18 (Schmidt & Hunter, 1998), and the predictive value of experience is concentrated in the early years on a job rather than accumulating indefinitely (Schmidt & Hunter, 1998). A résumé is, for the most part, a list of exactly these proxies. The literature’s verdict is not that résumés are worthless — it is that a funnel which filters hard on proxies before ever measuring job-relevant knowledge or behaviour discards most of the predictive power available to it.

Applied at Future Proof

How Future Proof™ applies this.

Hiring campaigns on the platform lead with the top of the validity table instead of the bottom. Candidates start with subject batteries — job knowledge tests scoped to the specific role and stack — and work-relevant questions built from the role’s actual competencies, delivered and scored identically for every candidate. Résumé fields don’t gate who gets measured: the structured assessment comes first, and comparable, job-relevant scores come out the other side.

See hiring campaigns

What the evidence doesn’t show

This literature is unusually deep, and precisely because it gets cited so confidently, its limits are worth stating plainly:

  • The corrections debate is not settled. The 2022 downward revision has itself been challenged — Oh, Le, and Roth argue that abandoning range-restriction corrections in concurrent studies overcorrects in the opposite direction (Oh et al., 2023). The exact magnitudes remain contested; the shift in rank order has proven more robust than any single number.
  • These are averages, not guarantees. A meta-analytic .42 for structured interviews describes well-built, consistently scored interviews. It says nothing about a loosely scripted panel that shares the label. Validity varies with job complexity and with the quality of execution.
  • The criterion is soft. Most validation studies use supervisor ratings as the performance measure. Ratings are themselves unreliable and partially contaminated, so “predicts job performance” mostly means “predicts how bosses rate performance” — related to, but not identical with, objective output.
  • Validity is not a complete hiring policy. Predictor choice interacts with subgroup score differences, adverse impact, cost, and candidate experience. A validity ranking is one input to selection-system design, not the whole design.
  • Generalization has edges. The evidence base is dominated by Western labour markets, incumbent samples, and typical rather than maximal performance. Effects in other contexts are plausible but less directly established.

None of these caveats rescues the bottom of the table. Under every correction regime yet proposed, structured, job-relevant measurement out-predicts credential proxies. The century-long argument is now about how large the gap is — not about which direction it points.

References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

  1. Schmidt, F.L., & Hunter, J.E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin 124(2): 262–274. DOI
  2. Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOIPDF
  3. Hunter, J.E., & Hunter, R.F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin 96(1): 72–98. DOI
  4. McDaniel, M.A., Whetzel, D.L., Schmidt, F.L., & Maurer, S.D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology 79(4): 599–616. DOI
  5. Huffcutt, A.I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology 79(2): 184–190. DOI
  6. Barrick, M.R., & Mount, M.K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology 44(1): 1–26. DOI
  7. Roth, P.L., Bobko, P., & McFarland, L.A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology 58(4): 1009–1037. PDF
  8. Dana, J., Dawes, R., & Peterson, N. (2013). Belief in the unstructured interview: The persistence of an illusion. Judgment and Decision Making 8(5): 512–520. PDF
  9. Kuncel, N.R., Klieger, D.M., Connelly, B.S., & Ones, D.S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology 98(6): 1060–1072. PDF
  10. Oh, I.-S., Le, H., & Roth, P.L. (2023). Revisiting Sackett et al.’s (2022) rationale behind their recommendation against correcting for range restriction in concurrent validation studies. Journal of Applied Psychology 108. PDF
Assessment science in production

Run a hiring funnel that measures what predicts.

Book a 20-minute demo and we’ll build a campaign for one of your real roles — subject battery, work-relevant questions, structured scoring — and show you what the funnel looks like when assessment comes before the résumé sort.

10 citations Reviewed July 2026 Open peer review welcomed