Situational judgment: testing the choices before the job.
Present a realistic dilemma, offer four plausible responses, score the choice against expert judgment. Ninety years of evidence says this humble format predicts job performance, adds signal beyond ability and personality, and travels with fewer adverse side effects — under conditions worth knowing. How Future Proof™ builds scenario items.
The finding: Situational judgment tests (SJTs) — scenario-plus-response-options items scored against an expert key — predict job performance with meta-analytic validity in the mid-.20s to low-.30s, add incremental validity over cognitive ability and personality, and show smaller subgroup differences than ability tests. Longitudinal work finds interpersonal SJT scores predicting performance years into a career.
The mechanism: SJTs measure procedural knowledge about effective behavior — knowing what one should do in consequential, ambiguous situations. That knowledge sits closer to the job than abstract traits, which is why it adds signal the trait measures miss.
The product: Future Proof’s assessment suite drafts role-specific scenarios from real critical incidents, keeps scoring keys under expert control, and uses knowledge-focused instructions — the variant the evidence favors for resistance to faking.
In this article
- 01Anatomy of an item
- 02One item, taken apart
- 03What the validity evidence says
- 04What an SJT actually measures
- 05The fairness case, spelled out
- 06Format matters: the video finding
- 07Building a key that deserves the name
- 08What the evidence doesn’t show
- 09What this means for practice
A customer-success lead discovers, two days before renewal, that a major client has quietly been evaluating a competitor. Four responses are on the table: escalate to leadership, call the client directly, prepare a counter-proposal first, or loop in the account’s executive sponsor. None is absurd. The candidate’s ranking of those options — against the ranking of people who do this job well — is a measurement. That is a situational judgment test. It is one of the oldest formats in continuous use in personnel assessment, with roots in the 1940s and forerunners older still (Motowidlo, Dunnette & Carter, 1990).
The format’s modern revival began when Motowidlo and colleagues reframed it honestly. An SJT is not a simulation of behavior but a low-fidelity simulation — a paper-and-pencil stand-in for the judgment the job demands. What has piled up since is a validity literature large enough for meta-analysis several times over. It is also specific enough to say what SJTs measure, what they add, and where they break.
Anatomy of an item
A situational judgment item has four working parts, and each is a place where quality is won or lost. The stem is the scenario itself: a situation with real stakes, incomplete information, and no clearly correct move. (A scenario with an obvious answer measures reading, not judgment.) The best stems descend from the critical-incident tradition. They are cleaned-up versions of situations that actually happened, gathered from people who do the job. That is what anchors the test to the role rather than to a test-writer’s imagination (Motowidlo et al., 1990).
The response options are harder to write than the stem. Every option must be something a reasonable candidate might genuinely do. The moment one option reads as a caricature, the item collapses into multiple-choice etiquette. The instructions decide what the item measures: “what should you do” versus “what would you do” — a one-word difference with measurable effects we return to below. And the scoring key turns choices into numbers, usually by scoring agreement with the pooled judgment of experienced performers.
Each part fails in its own way. Vague stems measure nothing. Implausible options leak the answer. Would-do instructions invite theater. A key built from one manager’s opinions measures closeness to that manager. The literature’s quality bar is, in effect, a checklist against these four failures (Campion et al., 2014).
One item, taken apart
Here is the anatomy assembled, in an item a customer-success team might field. Stem: “A strategic account’s day-to-day contact tells you, informally, that your platform lost an internal comparison and the renewal is at risk. Your executive sponsor at the account is traveling and unreachable for three days. The renewal call is in two weeks. What should you do first?”
The options: (a) escalate at once to your own leadership and request an executive-to-executive call. (b) Thank the contact, ask what the comparison measured, and request the evaluation criteria. (c) Begin building a counter-proposal aimed at the likely gaps. (d) Wait for the sponsor’s return before acting, to avoid signaling alarm.
Every option is defensible — that is the design requirement doing its work. The key, built from experienced account leaders’ pooled judgment, typically rewards (b) most. It turns a rumor into usable intelligence before any move is committed, and it costs nothing. Option (c) optimizes early against unknown criteria; (a) spends escalation capital on an unverified report; (d) mistakes inaction for caution while the clock runs. A candidate’s ranking across a dozen such scenarios measures exactly the procedural knowledge the construct research describes — knowing which moves the situation rewards. No résumé line or self-description retrieves it as directly.
What the validity evidence says
The first full meta-analysis put SJT validity for predicting job performance at roughly ρ ≈ .34 (McDaniel, Morgeson, Finnegan, Campion & Braverman, 2001). Two decades of stricter corrections later, the landmark re-estimate of all selection procedures places SJTs around .26. That is the same band as cognitive ability under the new corrections, below structured interviews and job-knowledge tests, and well above unstructured interviews (Sackett, Zhang, Berry & Lievens, 2022). Two more properties do much of the practical work. SJTs add incremental validity over cognitive ability and Big Five measures — they capture signal those tools miss (McDaniel et al., 2001), (McDaniel, Hartman, Whetzel & Grubb, 2007). And their subgroup differences are smaller than those of ability tests — which matters for fairness and for legal exposure alike (Whetzel, McDaniel & Nguyen, 2008).
≈ .26 Where two decades of stricter corrections place SJT operational validity — the same band as cognitive ability under the new corrections, with incremental validity over ability and personality on top (Sackett, Zhang, Berry & Lievens, 2022).
The most striking single result in the literature follows people over time. An interpersonal SJT given to medical school applicants predicted not just grades but actual job performance as physicians years later. It added validity beyond the knowledge-heavy admission tests it sat beside (Lievens & Sackett, 2012), (Lievens & Patterson, 2011). Judgment measured before the career predicted the career.
What an SJT actually measures
For decades the honest answer was “it depends on the test.” Construct-mapping work brought order. Most SJTs are soaked in leadership and interpersonal content, plus a general factor researchers came to read as procedural knowledge about effective behavior — in plainer terms, knowing what one should do (Christian, Edwards & Bradley, 2010). The theoretical capstone reframes SJTs as measures of general domain knowledge: unspoken rules about when certain behaviors pay (when to confront, when to defer, when to escalate). People acquire these rules unevenly, and jobs reward them (Lievens & Motowidlo, 2016).
One design choice changes the measurement more than any other: the response instructions. Ask “what should you do?” and the test behaves like a knowledge measure — more strongly tied to cognitive ability, and markedly harder to fake. Ask “what would you do?” and it behaves like a personality measure — easier to fake, easier to coach (McDaniel et al., 2007). When candidates have reason to game the test, the evidence favors the knowledge form.
One word changes the construct. “What should you do?” makes the test behave like a knowledge measure — markedly harder to fake; “what would you do?” makes it behave like personality — more fakable, more coachable. A vendor who cannot tell you which instruction their SJT uses has not read their own instrument.
The general-domain-knowledge reading also explains the format’s most useful practical property. SJT knowledge is learnable, and knowing what to do is a real, separable step on the way to doing it. A new team lead who cannot yet run a hard conversation smoothly can still know that public criticism of a struggling report is the wrong opening move. And the candidate who lacks even that knowledge is a different hiring bet from the one who has it but needs practice.
This is why the same scenario bank does double duty across the employee lifecycle. With a scoring key, it is a selection instrument. With feedback attached to each option, it becomes situated training — teaching the judgment it was built to measure. The construct that once looked like a psychometric embarrassment — “it measures knowing what one should do” — turns out to be exactly the thing a company can act on.
Knowing what one should do — measured before anyone is trusted to do it.The construct the SJT literature converged on: procedural knowledge about effective behavior (Christian et al., 2010; Lievens & Motowidlo, 2016).
The fairness case, spelled out
The subgroup finding deserves more than its one sentence, because it is where SJTs earn a seat that validity alone would not give them. Cognitive ability tests, for all their predictive power, produce some of the largest group gaps in the selection toolkit. Companies that lean on them alone buy their validity at a measurable adverse-impact cost. SJTs sit at a different point on that frontier: meaningful validity with much smaller subgroup gaps (Whetzel, McDaniel & Nguyen, 2008) — and smaller still in video form, once the reading toll is removed (Chan & Schmitt, 1997). So a battery that blends ability, knowledge, structured interviews, and situational judgment can be tuned to hold validity while shrinking impact. That is precisely the trade a defensible modern process is asked to make.
Candidates, for their part, consistently rate scenario-based assessment among the most acceptable formats they meet. It looks like the job, which abstract puzzles conspicuously do not. That perception is not mere comfort. A test that previews real situations also informs the candidate — and some will conclude the role is not for them, and select out before anyone spends an interview loop finding out the hard way. An instrument that predicts performance, shrinks impact, and improves the candidate’s information at the same time is doing three jobs for one delivery cost. That is the quiet reason the format has survived every fashion cycle since the 1940s.
Format matters: the video finding
SJTs need not be walls of text. In a direct comparison, a video-based SJT showed higher validity for interpersonal criteria and much smaller subgroup gaps than the same content in writing (Chan & Schmitt, 1997). Much of a written SJT’s adverse impact turns out to be a reading toll unrelated to the judgment being measured. Richer media also buys face validity: candidates treat realistic scenarios as a preview of the job, not an arbitrary hurdle. The state-of-the-field review counts these design levers — media, instructions, key construction — as the difference between SJTs that earn their validity and SJTs that merely resemble them (Campion, Ployhart & MacKenzie, 2014).
Building a key that deserves the name
Everything in an SJT rests, in the end, on the claim that its key encodes real expertise. So the key-building step deserves more scrutiny than it usually gets. The standard approach pools ratings from experienced, well-regarded performers and scores candidates by agreement with the pool. Before trusting it, check the pool itself. If your experts disagree sharply about an item, the item has no key — it has a controversy, and a candidate should not be scored on which side of an internal debate they land. Items where expert agreement is weak get rewritten or cut.
It is also worth auditing what the surviving key rewards. Keys built casually tend to encode company folklore — always escalate, never push back — rather than effectiveness. A folklore key selects for conformity with yesterday’s habits.
Day to day, two more habits keep a deployed SJT honest. Items age. Scenarios reference tools, structures, and norms that drift, and a bank that is never refreshed slowly detaches from the job it samples — on top of the exposure problem, since any static, high-stakes item pool eventually leaks. And validity is local until checked. The meta-analytic averages license the method, not your copy of it, so track scores against later performance data where volume permits. These are the same habits every serious assessment program runs; the SJT’s only quirk is that its content looks so conversational that teams forget it is a psychometric instrument that needs upkeep.
What the evidence doesn’t show
- SJTs are not interchangeable. Validity generalizes across well-built tests, but “well-built” is load-bearing: scenarios must come from real critical incidents, and scoring keys from genuine expert consensus. A plausible-sounding scenario with an armchair key measures agreement with its author, nothing more (Campion et al., 2014).
- Behavioral-tendency versions are coachable. “Would do” instructions invite impression management, and retest and coaching effects are real. High-stakes use should prefer knowledge instructions and refreshable item pools (McDaniel et al., 2007).
- Cross-cultural keys don’t always travel. What counts as the effective response — how directly to confront, when to involve a superior — varies across cultures; keys built in one context can quietly punish candidates operating under different, equally valid norms.
- Mid-table validity means mid-table use. An SJT is a component of a battery, not a replacement for one. The evidence supports it precisely because it adds unique signal beside ability, knowledge, and structured interviews — not instead of them (Sackett et al., 2022).
Where the evidence stops
- 1SJTs are not interchangeable
- 2Behavioral-tendency versions are coachable
- 3Cross-cultural keys don’t always travel
- 4Mid-table validity means mid-table use
What this means for practice
Source scenarios from the job, not from imagination. Collect critical incidents from the people doing the work now, and let the situations candidates will actually face define the test. The collection step is lighter than teams expect. A handful of structured conversations per role yields more usable stems than a month of test-writer invention — and the resulting items carry a face validity no generic bank can match. Build keys from expert consensus; check that the experts agree with each other before assuming a candidate should agree with them; audit what the key rewards. Use should-do instructions where stakes are high, and keep the pool fresh.
Deliver scenarios in rich media where reading load would otherwise pollute the measure — the video finding is thirty years old and still routinely ignored. Reuse the investment: the same validated scenarios, with feedback attached to each option, become the most job-relevant training content a company owns. And place SJT scores where the evidence places them — one calibrated signal in a structured battery, weighed beside ability, knowledge, and structured interviews, not crowning or replacing them. A method that has survived ninety years of psychometric scrutiny deserves exactly that seat at the table: earned, specific, and not the head of it.
How Future Proof™ applies this: scenario items with expert keys.
The assessment suite treats SJTs the way the literature prescribes. Scenario drafts are generated from role-specific critical incidents — the situations your own teams report — then reviewed and keyed by your subject-matter experts, whose consensus is checked before any item goes live. Items use knowledge-focused instructions, rotate from refreshable pools, and report into the same structured scorecard as ability, knowledge, and interview evidence, weighted as one signal among several. Judgment, measured before the job — and never by an armchair key.
See the assessment suite →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
The evidence, by year
- 1990Motowidlo
- 1997Chan
- 2001McDaniel
- 2007McDaniel
- 2008Whetzel
- 2010Christian
- 2011Lievens
- 2012Lievens
- 2014Campion
- 2016Lievens
- 2022Sackett
- Motowidlo, S.J., Dunnette, M.D., & Carter, G.W. (1990). An alternative selection procedure: The low-fidelity simulation. Journal of Applied Psychology 75(6): 640–647. PDF
- McDaniel, M.A., Morgeson, F.P., Finnegan, E.B., Campion, M.A., & Braverman, E.P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology 86(4): 730–740. DOI
- Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOI
- McDaniel, M.A., Hartman, N.S., Whetzel, D.L., & Grubb, W.L. (2007). Situational judgment tests, response instructions, and validity: A meta-analysis. Personnel Psychology 60(1): 63–91. PDF
- Whetzel, D.L., McDaniel, M.A., & Nguyen, N.T. (2008). Subgroup differences in situational judgment test performance: A meta-analysis. Human Performance 21(3): 291–309. PDF
- Lievens, F., & Sackett, P.R. (2012). The validity of interpersonal skills assessment via situational judgment tests for predicting academic success and job performance. Journal of Applied Psychology 97(2): 460–468. PDF
- Lievens, F., & Patterson, F. (2011). The validity and incremental validity of knowledge tests, low-fidelity simulations, and high-fidelity simulations for predicting job performance in advanced-level high-stakes selection. Journal of Applied Psychology 96(5): 927–940. PDF
- Christian, M.S., Edwards, B.D., & Bradley, J.C. (2010). Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology 63(1): 83–117. PDF
- Lievens, F., & Motowidlo, S.J. (2016). Situational judgment tests: From measures of situational judgment to measures of general domain knowledge. Industrial and Organizational Psychology 9(1): 3–22. PDF
- Chan, D., & Schmitt, N. (1997). Video-based versus paper-and-pencil method of assessment in situational judgment tests: Subgroup differences in test performance and face validity perceptions. Journal of Applied Psychology 82(1): 143–159. PDF
- Campion, M.C., Ployhart, R.E., & MacKenzie, W.I. (2014). The state of research on situational judgment tests: A content analysis and directions for future research. Human Performance 27(4): 283–310. PDF
Test the judgment your roles actually require.
Book a 20-minute demo. Bring three real situations your team faced this quarter — we’ll show you scenario items drafted from them, keyed by your experts, scored in a structured battery.