The Big Five at work: evidence and limits.
Personality testing gets dismissed as corporate astrology and sold as a crystal ball. Three decades of meta-analyses support neither. What the Big Five actually predicts at work, what the 2022 recalibration changed, and why Future Proof™ treats personality as one signal among several.
The finding: Across three decades of meta-analyses, conscientiousness predicts job performance in essentially every job family — with corrected correlations in the low .20s, revised to roughly .19 in 2022. The other four traits predict only where the work demands them. The effect is real, replicated, and modest.
The mechanism: Traits are stable behavioral tendencies. Conscientious people set goals, persist, and follow through, and those small daily advantages accumulate across almost any job’s tasks. For the other traits, prediction appears only when the tendency matches what the role actually requires — sociability in sales, for instance, but not in solo technical work.
The product: Future Proof uses a 120-item Big Five instrument with role-fit profiles as one signal among several — feeding structured interviews and development plans, never acting as a sole screen.
In this article
- 01From pessimism to a workable taxonomy
- 02What the follow-up meta-analyses converged on
- 03The development use case, held to the same standard
- 04Facets, not just factors
- 05The 2022 recalibration
- 06What the evidence doesn’t show
- 07Using personality signals responsibly
Two conversations about personality testing run in parallel and never meet. In one, practitioners give tests, read profiles, and make decisions. In the other, researchers publish meta-analyses — studies that pool many studies — whose effect sizes would startle most of the people using the tools. This article is the meeting: the actual numbers, their real uses, and the specific claims that exceed them.
Personality testing holds a strange position in workplace decisions. It is one of the most heavily meta-analyzed topics in work psychology — and, at the same time, one of the most casually oversold categories in HR technology. Somewhere between “personality tests are corporate astrology” and “our test finds your future top performers” sits a real body of evidence. Its findings have survived three decades of replication. Its boundary conditions are the part most vendor decks quietly omit.
This article walks through that evidence by way of three landmarks. First, the 1991 meta-analysis that made personality respectable in selection research again (Barrick & Mount, 1991). Second, the 2013 work showing that the five broad factors are coarser than the data beneath them (Judge et al., 2013). Third, the 2022 re-analysis that revised the whole field’s effect sizes (Sackett et al., 2022). Then the limits, stated as plainly as the findings.
From pessimism to a workable taxonomy
The field’s redemption arc is worth telling in full, because it explains both the evidence’s strength and its precise shape. A literature rebuilt on a taxonomy after decades in the wilderness carries its structure visibly. For much of the twentieth century, the considered scientific position was that personality testing had no place in hiring. Reviewing the scattered evidence in 1965, Robert Guion and Richard Gottier concluded that personality measures could not, in good conscience, be recommended as a basis for hiring decisions (Guion & Gottier, 1965). The problem was not only weak results. It was chaos: hundreds of scales with clashing names, measuring overlapping constructs, with no shared frame for adding up what any of them found.
The Big Five changed the question. By the late 1980s, factor-analytic work had converged on five broad dimensions: conscientiousness, emotional stability, extraversion, agreeableness, and openness to experience. Nearly any personality scale could be sorted into them. A scattered literature became one that meta-analysis could actually sum up.
Barrick and Mount’s 1991 meta-analysis in Personnel Psychology was the turning point (Barrick & Mount, 1991). They sorted 117 validity studies into the new taxonomy. Conscientiousness predicted job performance in every job group they examined — professionals, police, managers, sales, skilled and semi-skilled work — with a corrected validity in the low .20s. Extraversion predicted performance in jobs that run on dealing with people, chiefly sales and management. Openness and extraversion predicted how well people did in training. The rest of the links were small and tied to context.
The same year, Tett, Jackson, and Rothstein added a result that still shapes good practice: validity depends on how traits are chosen. Studies that picked traits from a job analysis — on a hypothesis about what the work demands — produced far higher validities than fishing designs that correlated everything with everything (Tett et al., 1991).
It is difficult in the face of this summary to advocate, with a clear conscience, the use of personality measures in most situations as a basis for making employment decisions about people.Guion & Gottier 1965 — the consensus the Big Five era overturned
What the follow-up meta-analyses converged on
One landmark result invites the question every replication era has taught the field to ask. Does it survive tougher conditions — cleaner tests, independent teams, summaries of summaries? The decade after 1991 was largely that stress test, and the core result held. Hurtz and Donovan cut the evidence base down to tests explicitly built to measure the Big Five, rather than older scales retrofitted into the taxonomy. They found the same pattern at slightly lower magnitudes (Hurtz & Donovan, 2000). Conscientiousness was again the most consistent predictor, with emotional stability behind it.
Barrick, Mount, and Judge then pooled fifteen prior meta-analyses into a second-order summary. Conscientiousness held up across jobs and outcome measures. Emotional stability generalized more weakly, and the other three traits predicted only for certain jobs or certain outcomes (Barrick et al., 2001).
The generality claim deserves its precision. “Predicts across occupations” means the conscientiousness coefficient stayed positive and meaningful in every job family examined. No other trait matched that reach, and few predictors of any kind do. The other four traits are not weaker so much as conditional. That is itself usable design input: their profiles belong in role-fit logic, not in universal screens.
The pattern also extends beyond task performance. In leadership research, extraversion emerged as the steadiest trait correlate of who gets seen as a leader, and of who is rated effective in the role (Judge et al., 2002). And outside work entirely, long-run evidence shows personality traits predicting mortality, divorce, and career attainment with effects on par with social class and cognitive ability (Roberts et al., 2007).
It is worth pausing on magnitude, because this is where marketing and evidence part ways. A corrected correlation in the low .20s explains a few percent of the variance in job performance. That is genuinely modest. It is also genuinely useful: the signal is cheap to measure, points the same direction across nearly all jobs, and carries information that ability tests do not. Honest use of personality data requires holding both halves of that sentence at once.
The development use case, held to the same standard
Selection is where personality data faces its hardest test: adversarial conditions, faking incentives, weighty decisions. Much of its defensible value lies in the gentler use case — development. Traits are stable dispositions with real long-term reach (Roberts et al., 2007). That makes an honest profile useful the way a map of one’s own tendencies is useful.
Think of the low-conscientiousness contributor who builds outside scaffolding for follow-through. Or the introverted new manager who schedules recovery around the role’s social load. Or the team that learns why two members keep colliding over process. None of this requires the profile to predict performance. It requires only that the profile be measurement rather than horoscope.
The same standard rules out the development market’s favorite failure mode: typologies. Sorting people into named types — colors, letters, animals — throws away the dimensional structure the Big Five evidence rests on. It imposes category borders the trait data do not contain, and it trades test-retest stability for being easy to remember.
A type-based workshop may still be useful as a conversation starter. But its scientific pedigree is not this literature’s. Firms citing the Big Five’s evidence while running typologies are borrowing credibility across a real divide. The evidence-based development tool is the same one the selection evidence prefers: dimensional, facet-level, and honest about its error bands.
Facets, not just factors
Everything to this point has treated the five factors as the natural units of study, and for meta-analytic housekeeping they were essential. But the factors were always a compression choice, and compression throws information away. That fact shapes what a test must look like to earn its claims. The five factors are summaries, and summaries blur. Conscientiousness, for example, bundles achievement striving together with orderliness and caution. Those tendencies are correlated but not interchangeable — and they plausibly matter differently for a founder, an auditor, and an air-traffic controller.
Judge, Rodell, Klinger, Simon, and Crawford tested this systematically. They compared layers of the five-factor model — broad domains, middle aspects, narrow facets — as predictors of job performance. Narrower traits carried predictive information that the broad domain scores averaged away. The size of the gain depended on the trait and the outcome in question (Judge et al., 2013).
The practical consequence is unglamorous but important: measurement resolution costs items. A ten-item screener can estimate five broad scores, noisily. Resolving the facets beneath them — the level where much of the predictive signal lives — takes tests an order of magnitude longer. That is why serious Big Five inventories run to a hundred items or more.
The 2022 recalibration
In 2022, Sackett, Zhang, Berry, and Lievens re-examined the meta-analytic base of the selection literature. They argued that the standard statistical corrections for range restriction had been too aggressive across the board (Sackett et al., 2022). Their revised estimates moved almost every predictor downward. Structured interviews rose to the top of the rankings at a corrected validity of roughly .42. General cognitive ability fell from its canonical .51 to roughly .31; conscientiousness landed near .19.
≈.19 Conscientiousness’s corrected validity after the 2022 audit built to deflate the field’s numbers — essentially where Barrick and Mount put it thirty years earlier, because it was never propped up by the corrections that flattered ability tests (Sackett et al., 2022).
Note what surviving a hostile audit means for the evidence. The 2022 project was built to deflate the field’s numbers — and it deflated nearly everything except the claim this article centers on. Conscientiousness’s estimate had never been inflated by the aggressive corrections that flattered ability tests. The trait’s modest coefficient turned out to be the durable kind of modest.
There are two honest readings. In absolute terms, personality validities remain modest — close to where Barrick and Mount put them thirty years earlier, which is itself a kind of replication. In relative terms, personality’s standing improved, because the predictors it competes with fell further. The deeper lesson of the re-analysis, though, concerns combination. No single predictor dominates, and the defensible strategy is a battery of signals anchored by structured methods rather than any one score (Sackett et al., 2022).
What the evidence doesn’t show
The literature’s honesty about its own limits is its best credential, and the limits are load-bearing for anyone buying or deploying tests. Four boundaries, stated plainly — each one refutes a claim now circulating in some vendor’s deck.
It does not show that personality can carry a hiring decision. A validity near .19 leaves most of the variance in performance unexplained. A vendor who says a personality profile by itself “identifies top performers” is claiming something no meta-analysis in this literature has found.
It has not settled the faking problem. Big Five tests are self-reports, and applicants have obvious incentives to answer strategically. In a pointed 2007 exchange, Morgeson and colleagues — several of them former editors of the field’s flagship journals — argued for real caution about personality tests in selection. Observed validities were low, they held, and applicant faking too stubborn to fix (Morgeson et al., 2007). Ones and colleagues replied that corrected validities are meaningful, that faking weakens but does not erase predictive power, and that personality measures show far smaller group differences than cognitive ability tests (Ones et al., 2007). Both papers repay reading; the disagreement was never fully resolved.
Big Five instruments are self-reports, applicants have obvious incentives to answer strategically, and the field’s own former journal editors called the faking problem serious enough to warrant real caution in selection (Morgeson et al., 2007). The defensible deployment keeps personality as one input to structured human decisions — never an automated screen.
It does not show that more of a trait is always better. There is evidence that some trait–performance relationships are curvilinear: gains flatten, and in some jobs reverse, at high levels of conscientiousness and emotional stability (Le et al., 2011). Very high conscientiousness can shade into rigidity and perfectionism. Profile logic — is this level of this trait suited to this work — has better support than more-is-better logic.
It does not measure ability. The Big Five describes typical tendencies — how a person is disposed to behave across situations — not skills, knowledge, or maximal performance. A highly conscientious novice is still a novice. Personality data speaks to how someone will tend to work, never to whether they can do the work, which is why it can complement skills evidence but cannot substitute for it.
How Future Proof™ applies this — as one signal, never a verdict.
Future Proof’s assessment suite includes a 120-item Big Five instrument — long enough to resolve facet-level signal, which is where the hierarchical evidence says much of the information lives. Role-fit profiles map traits to specific roles the way the confirmatory literature prescribes: the profile is defined from the role’s demands before anyone is scored. The output sits alongside skills diagnostics and structured interview rubrics in a candidate or learner profile. It is never used as a sole screen, and it never auto-rejects anyone.
See the assessment suite →Using personality signals responsibly
Held together — the modest validities, the facet resolution, the faking debate, the curves that bend at the top — the literature does prescribe. Its prescriptions are specific enough to audit any deployment against. Fifty years compress into four practices that separate defensible use from pseudoscience:
- Choose traits from the job, not from the dashboard. Confirmatory, job-analysis-driven trait selection produced markedly higher validities than exploratory correlation-hunting (Tett et al., 1991). A role profile should exist before the first candidate is scored.
- Measure below the domain. Facet-level assessment preserves signal that broad-factor scores blur (Judge et al., 2013) — which requires instruments long enough to resolve facets reliably.
- Combine, don’t crown. The best-validated selection procedures are structured methods; personality earns its place as an increment to them, not a replacement for them (Sackett et al., 2022).
- Keep consequential decisions with humans. Given the unresolved faking debate and the modest absolute validities, the defensible role for personality data is informing structured interviews, development plans, and team conversations — not automated rejection.
Where the evidence stops
- 1Choose traits from the job, not from the dashboard
- 2Measure below the domain
- 3Combine, don’t crown
- 4Keep consequential decisions with humans
The Big Five literature is one of applied psychology’s quiet success stories precisely because it is honest about its own size. Its central finding survived a re-analysis designed to shrink it. That is what a real effect looks like: modest, replicated, and useful — provided nobody asks it to do a job it never claimed it could do.
Selected papers.
These are the studies cited above — the meta-analytic backbone of personality-at-work research, not an exhaustive bibliography.
The evidence, by year
- 1965Guion
- 1991Barrick
- 1991Tett
- 2000Hurtz
- 2001Barrick
- 2002Judge
- 2007Roberts
- 2007Morgeson
- 2007Ones
- 2011Le
- 2013Judge
- 2022Sackett
- Guion, R.M., & Gottier, R.F. (1965). Validity of personality measures in personnel selection. Personnel Psychology 18(2): 135–164. PDF
- Tett, R.P., Jackson, D.N., & Rothstein, M. (1991). Personality measures as predictors of job performance: A meta-analytic review. Personnel Psychology 44(4): 703–742. PDF
- Hurtz, G.M., & Donovan, J.J. (2000). Personality and job performance: The Big Five revisited. Journal of Applied Psychology 85(6): 869–879. DOI
- Barrick, M.R., Mount, M.K., & Judge, T.A. (2001). Personality and performance at the beginning of the new millennium: What do we know and where do we go next? International Journal of Selection and Assessment 9(1–2): 9–30. PDF
- Judge, T.A., Bono, J.E., Ilies, R., & Gerhardt, M.W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology 87(4): 765–780. DOI
- Roberts, B.W., Kuncel, N.R., Shiner, R., Caspi, A., & Goldberg, L.R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science 2(4): 313–345. PDF
- Judge, T.A., Rodell, J.B., Klinger, R.L., Simon, L.S., & Crawford, E.R. (2013). Hierarchical representations of the five-factor model of personality in predicting job performance: Integrating three organizing frameworks with two theoretical perspectives. Journal of Applied Psychology 98(6): 875–925. PDF
- Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOI
- Morgeson, F.P., Campion, M.A., Dipboye, R.L., Hollenbeck, J.R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology 60(3): 683–729. PDF
- Ones, D.S., Dilchert, S., Viswesvaran, C., & Judge, T.A. (2007). In support of personality assessment in organizational settings. Personnel Psychology 60(4): 995–1027. PDF
- Le, H., Oh, I.-S., Robbins, S.B., Ilies, R., Holland, E., & Westrick, P. (2011). Too much of a good thing: Curvilinear relationships between personality traits and job performance. Journal of Applied Psychology 96(1): 113–133. PDF
Personality is one signal. See the full picture.
Book a 20-minute demo and see how a role-fit profile sits alongside skills diagnostics and structured interview rubrics in Future Proof — and why no single score ever decides an outcome.