The Big Five at work: evidence and limits.
Personality testing gets dismissed as corporate astrology and sold as a crystal ball. Three decades of meta-analyses support neither. What the Big Five actually predicts at work, what the 2022 recalibration changed, and why Future Proof™ treats personality as one signal among several.
The finding: Across three decades of meta-analyses, conscientiousness predicts job performance in essentially every job family — with corrected correlations in the low .20s, revised to roughly .19 in 2022. The other four traits predict only where the work demands them. The effect is real, replicated, and modest.
The mechanism: Traits are stable behavioral tendencies. Conscientious people set goals, persist, and follow through, and those small daily advantages accumulate across almost any job’s tasks. For the other traits, prediction appears only when the tendency matches what the role actually requires — sociability in sales, for instance, but not in solo technical work.
The product: Future Proof uses a 120-item Big Five instrument with role-fit profiles as one signal among several — feeding structured interviews and development plans, never acting as a sole screen.
Personality assessment occupies a strange position in workplace decision-making. It is one of the most heavily meta-analyzed topics in industrial-organizational psychology — and, at the same time, one of the most casually oversold categories in HR technology. Somewhere between “personality tests are corporate astrology” and “our assessment identifies your future top performers” sits a real body of evidence, with findings that have survived three decades of replication and boundary conditions that most vendor decks quietly omit.
This article walks through that evidence by way of three landmarks: the 1991 meta-analysis that made personality respectable in selection research again (Barrick & Mount, 1991), the 2013 work showing that the five broad factors are coarser than the data beneath them (Judge et al., 2013), and the 2022 re-analysis that revised the entire field’s effect sizes (Sackett et al., 2022). Then the limits, stated as plainly as the findings.
From pessimism to a workable taxonomy
For much of the twentieth century, the considered scientific position was that personality testing had no place in hiring. Reviewing the scattered evidence in 1965, Robert Guion and Richard Gottier concluded that personality measures could not, in good conscience, be recommended as a basis for employment decisions (Guion & Gottier, 1965). The problem was not only weak results. It was incoherence: hundreds of scales with incompatible names, measuring overlapping constructs, with no shared framework for aggregating what any of them found.
The Big Five changed the question. By the late 1980s, factor-analytic work had converged on five broad dimensions — conscientiousness, emotional stability, extraversion, agreeableness, and openness to experience — into which nearly any personality scale could be sorted. A fragmented literature became one that meta-analysis could actually summarize.
Barrick and Mount’s 1991 meta-analysis in Personnel Psychology was the turning point (Barrick & Mount, 1991). Sorting 117 validity studies into the new taxonomy, they found that conscientiousness predicted job performance across every occupational group they examined — professionals, police, managers, sales, skilled and semi-skilled work — with a corrected validity in the low .20s. Extraversion predicted performance in jobs that run on interpersonal interaction, chiefly sales and management. Openness and extraversion predicted training proficiency. The remaining relationships were small and situation-specific.
The same year, Tett, Jackson, and Rothstein added a result that still shapes good practice: validity depends on how traits are chosen. Studies that selected traits confirmatorily — from a job analysis, on a hypothesis about what the work demands — produced substantially higher validities than exploratory designs that correlated everything with everything (Tett et al., 1991).
It is difficult in the face of this summary to advocate, with a clear conscience, the use of personality measures in most situations as a basis for making employment decisions about people.Guion & Gottier 1965 — the consensus the Big Five era overturned
What the follow-up meta-analyses converged on
The decade after 1991 was largely a stress test, and the core result held. Hurtz and Donovan restricted the evidence base to instruments explicitly built to measure the Big Five — rather than older scales retrofitted into the taxonomy — and found the same pattern at slightly lower magnitudes, with conscientiousness again the most consistent predictor and emotional stability behind it (Hurtz & Donovan, 2000). Barrick, Mount, and Judge then aggregated fifteen prior meta-analyses into a second-order summary: conscientiousness generalized across occupations and criteria, emotional stability generalized more weakly, and the other three traits predicted only for particular jobs or particular outcomes (Barrick et al., 2001).
The pattern also extends beyond task performance. In leadership research, extraversion emerged as the most consistent trait correlate of both who comes to be seen as a leader and who is rated effective in the role (Judge et al., 2002). And outside work entirely, longitudinal evidence shows personality traits predicting mortality, divorce, and occupational attainment with effects comparable in magnitude to socioeconomic status and cognitive ability (Roberts et al., 2007).
It is worth pausing on magnitude, because this is where marketing and evidence part ways. A corrected correlation in the low .20s explains a few percent of the variance in job performance. That is genuinely modest. It is also genuinely useful: the signal is cheap to measure, points the same direction across nearly all jobs, and carries information that ability tests do not. Honest use of personality data requires holding both halves of that sentence at once.
Facets, not just factors
The five factors are summaries, and summaries blur. Conscientiousness, for example, bundles achievement striving together with orderliness and cautiousness — tendencies that are correlated but not interchangeable, and that plausibly matter differently for a founder, an auditor, and an air-traffic controller. Judge, Rodell, Klinger, Simon, and Crawford tested this systematically, comparing hierarchical representations of the five-factor model — broad domains, intermediate aspects, narrow facets — as predictors of job performance. Narrower traits carried predictive information that the broad domain scores averaged away, with the size of the gain depending on the trait and the criterion in question (Judge et al., 2013).
The practical consequence is unglamorous but important: measurement resolution costs items. A ten-item screener can estimate five broad scores, noisily. Resolving the facets beneath them — the level where much of the predictive signal lives — requires instruments an order of magnitude longer, which is why serious Big Five inventories run to a hundred items or more.
The 2022 recalibration
In 2022, Sackett, Zhang, Berry, and Lievens re-examined the meta-analytic foundations of the personnel-selection literature and argued that the standard statistical corrections for range restriction had been systematically too aggressive (Sackett et al., 2022). Their revised estimates moved almost every predictor downward. Structured interviews rose to the top of the rankings at a corrected validity of roughly .42; general cognitive ability fell from its canonical .51 to roughly .31; conscientiousness landed near .19.
There are two honest readings. In absolute terms, personality validities remain modest — close to where Barrick and Mount put them thirty years earlier, which is itself a kind of replication. In relative terms, personality’s standing improved, because the predictors it competes with fell further. The deeper lesson of the re-analysis, though, concerns combination: no single predictor dominates, and the defensible strategy is a battery of signals anchored by structured methods rather than any one score (Sackett et al., 2022).
What the evidence doesn’t show
Four boundaries, stated plainly.
It does not show that personality can carry a hiring decision. A validity near .19 leaves the overwhelming majority of performance variance unexplained. A vendor claiming that a personality profile by itself “identifies top performers” is asserting something no meta-analysis in this literature has found.
It has not settled the faking problem. Big Five instruments are self-reports, and applicants have obvious incentives to answer strategically. In a pointed 2007 exchange, Morgeson and colleagues — several of them former editors of the field’s flagship journals — argued that observed validities were low and applicant faking intractable enough to warrant real caution about personality tests in selection (Morgeson et al., 2007). Ones and colleagues replied that corrected validities are meaningful, that faking attenuates but does not eliminate criterion-related validity, and that personality measures show far smaller subgroup differences than cognitive ability tests (Ones et al., 2007). Both papers repay reading; the disagreement was never fully resolved.
It does not show that more of a trait is always better. There is evidence that some trait–performance relationships are curvilinear: gains flatten, and in some jobs reverse, at high levels of conscientiousness and emotional stability (Le et al., 2011). Very high conscientiousness can shade into rigidity and perfectionism. Profile logic — is this level of this trait suited to this work — has better support than more-is-better logic.
It does not measure ability. The Big Five describes typical tendencies — how a person is disposed to behave across situations — not skills, knowledge, or maximal performance. A highly conscientious novice is still a novice. Personality data speaks to how someone will tend to work, never to whether they can do the work, which is why it can complement skills evidence but cannot substitute for it.
How Future Proof™ applies this — as one signal, never a verdict.
Future Proof’s assessment suite includes a 120-item Big Five instrument — long enough to resolve facet-level signal, which is where the hierarchical evidence says much of the information lives. Role-fit profiles map traits to specific roles the way the confirmatory literature prescribes: the profile is defined from the role’s demands before anyone is scored. The output sits alongside skills diagnostics and structured interview rubrics in a candidate or learner profile. It is never used as a sole screen, and it never auto-rejects anyone.
See the assessment suite →Using personality signals responsibly
Fifty years of this literature compress into four practices that separate defensible use from pseudoscience:
- Choose traits from the job, not from the dashboard. Confirmatory, job-analysis-driven trait selection produced markedly higher validities than exploratory correlation-hunting (Tett et al., 1991). A role profile should exist before the first candidate is scored.
- Measure below the domain. Facet-level assessment preserves signal that broad-factor scores blur (Judge et al., 2013) — which requires instruments long enough to resolve facets reliably.
- Combine, don’t crown. The best-validated selection procedures are structured methods; personality earns its place as an increment to them, not a replacement for them (Sackett et al., 2022).
- Keep consequential decisions with humans. Given the unresolved faking debate and the modest absolute validities, the defensible role for personality data is informing structured interviews, development plans, and team conversations — not automated rejection.
The Big Five literature is one of applied psychology’s quiet success stories precisely because it is honest about its own size. Its central finding survived a re-analysis designed to shrink it. That is what a real effect looks like: modest, replicated, and useful — provided nobody asks it to do a job it never claimed it could do.
Selected papers.
These are the studies cited above — the meta-analytic backbone of personality-at-work research, not an exhaustive bibliography.
-
Guion, R.M., & Gottier, R.F. (1965). Validity of personality measures in personnel selection. Personnel Psychology 18(2): 135–164. PDF
-
Tett, R.P., Jackson, D.N., & Rothstein, M. (1991). Personality measures as predictors of job performance: A meta-analytic review. Personnel Psychology 44(4): 703–742. PDF
-
Hurtz, G.M., & Donovan, J.J. (2000). Personality and job performance: The Big Five revisited. Journal of Applied Psychology 85(6): 869–879. DOI
-
Barrick, M.R., Mount, M.K., & Judge, T.A. (2001). Personality and performance at the beginning of the new millennium: What do we know and where do we go next? International Journal of Selection and Assessment 9(1–2): 9–30. PDF
-
Judge, T.A., Bono, J.E., Ilies, R., & Gerhardt, M.W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology 87(4): 765–780. DOI
-
Roberts, B.W., Kuncel, N.R., Shiner, R., Caspi, A., & Goldberg, L.R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science 2(4): 313–345. PDF
-
Judge, T.A., Rodell, J.B., Klinger, R.L., Simon, L.S., & Crawford, E.R. (2013). Hierarchical representations of the five-factor model of personality in predicting job performance: Integrating three organizing frameworks with two theoretical perspectives. Journal of Applied Psychology 98(6): 875–925. PDF
-
Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOI
-
Morgeson, F.P., Campion, M.A., Dipboye, R.L., Hollenbeck, J.R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology 60(3): 683–729. PDF
-
Ones, D.S., Dilchert, S., Viswesvaran, C., & Judge, T.A. (2007). In support of personality assessment in organizational settings. Personnel Psychology 60(4): 995–1027. PDF
-
Le, H., Oh, I.-S., Robbins, S.B., Ilies, R., Holland, E., & Westrick, P. (2011). Too much of a good thing: Curvilinear relationships between personality traits and job performance. Journal of Applied Psychology 96(1): 113–133. PDF
Personality is one signal. See the full picture.
Book a 20-minute demo and see how a role-fit profile sits alongside skills diagnostics and structured interview rubrics in Future Proof — and why no single score ever decides an outcome.