Structured interviews and BARS: the validity case.
Hiring managers trust their read of a candidate. Sixty years of selection research say the read is the problem. On the evidence for structured interviews and behaviorally anchored rating scales — the same architecture Future Proof™’s interview module scores against, with anchors, thresholds, and gates instead of gut feel.
The finding: Structured interviews carry roughly double the predictive validity of unstructured ones. The result has held from the first meta-analyses in the late 1980s through the most conservative re-analysis in 2022 — where the structured interview ranked as the single best predictor of job performance among common selection tools.
The mechanism: Structure removes noise from judgment. Every candidate gets the same job-derived questions; every answer is rated against behaviorally anchored scales; scores combine by rule. Anchors turn evaluation from impression into evidence-matching, and standardization makes candidates comparable — both raise reliability, which caps validity.
The product: Future Proof’s interview module scores every answer against per-competency BARS anchors with defined thresholds and gates — not gut impressions.
In this article
- 01What “structure” actually means
- 02The validity record
- 03Where BARS came from
- 04Why structure wins
- 05The defensibility case
- 06Keeping structure from decaying
- 07What the evidence doesn’t show
No hiring tool enjoys more confidence per unit of evidence than the free-flowing interview. None has been studied longer. The collision of the two — near-universal faith meeting six decades of measurement — produced the clearest before-and-after story in assessment science. The same activity, reorganized, roughly doubles its predictive power.
The job interview is the most widely used selection method in the world. And the version most companies use — an open conversation steered by the interviewer’s instincts — is the version the evidence supports least. That is not a new finding, and it is not a close call. Across six decades of meta-analytic work the pattern is stable. Interviews predict job performance well when they are structured, and poorly when they are not. The gap between the two is one of the largest method effects in personnel psychology.
The stakes grow with seniority. The bigger the role, the more likely the final gate is an unstructured chat with the most senior people in the building — the setup the evidence rates lowest, used where errors cost most. Why the unstructured interview persists is itself a research topic. Highhouse documented what he called practitioners’ stubborn reliance on intuition and subjectivity in selection — the belief that an experienced judge, freed from constraints, can read a candidate better than any protocol (Highhouse, 2008). Decision research points the other way. So, very specifically, does the interview literature.
What “structure” actually means
Before the evidence, the definitions. “Structured interview” gets claimed by everything from a fully anchored protocol to a manager with a printed question list. The validity numbers belong to specific setups, not to the label. Structure is also not a synonym for rigid or scripted. In the definitive modern review, Levashina, Hartwell, Morgeson, and Campion treat structure as any upgrade to the interview meant to improve its measurement by standardizing what is asked, how answers are judged, or both (Levashina et al., 2014).
An earlier review catalogued fifteen distinct components of structure and sorted them into two families (Campion, Palmer & Campion, 1997). Content structure governs the questions. Base them on a job analysis; ask every candidate the same ones; prefer behavioral and situational formats; limit ad-hoc prompting. Evaluation structure governs the judging. Rate each answer rather than forming one overall impression. Anchor the rating scale in described behaviors, take notes, use several interviewers, and combine scores by formula rather than by discussion.
Because structure is a bundle of parts, it is a dial rather than a switch. Huffcutt and Arthur sorted interview studies into four rising levels of structure — from no constraints on questions or scoring, up to fully standardized questions scored on anchored scales (Huffcutt & Arthur, 1994). That framing turned “do interviews work?” into a better question: how does validity move as structure increases?
The validity record
The structure question has now been answered so many times, by so many teams, under such different assumptions, that the record reads less like a finding than like a law being stress-tested. The answer arrived in waves of meta-analysis. Wiesner and Cronshaw found that structured interviews carried roughly twice the predictive validity of unstructured ones — corrected correlations of about .63 versus .20 (Wiesner & Cronshaw, 1988). Working from a larger database with more cautious assumptions, McDaniel, Whetzel, Schmidt, and Maurer found the same ordering at smaller values: about .44 for structured interviews against job-performance measures, versus about .33 for unstructured (McDaniel et al., 1994).
The level analysis filled in the shape of the curve. Validity climbed from roughly .20 at the lowest level of structure to roughly .57 at the highest. Most of the gain arrived by the second-highest level (Huffcutt & Arthur, 1994).
.63 vs .20 The corrected validity of structured versus unstructured interviews in the first major synthesis — roughly double the predictive power for the same interviewer hours (Wiesner & Cronshaw, 1988).
Summing up eighty-five years of selection research, Schmidt and Hunter put the structured interview at .51 — level with general mental ability among the strongest single predictors available — against .38 for the unstructured version (Schmidt & Hunter, 1998). Then Sackett, Zhang, Berry, and Lievens re-estimated the whole table, using more defensible corrections for range restriction. Nearly every predictor’s validity fell, but the ordering sharpened. The structured interview came out on top of the revised rankings, at roughly .42. The unstructured interview dropped far down the list (Sackett et al., 2022).
The exact coefficients move with the correction method. The direction never has.
The 2022 result deserves one more line of emphasis, because it flips the field’s old pecking order. For decades the cognitive ability test was selection’s crown jewel and the interview its tolerated ritual. Under the revised corrections, the best-run version of the ritual is the single strongest common predictor on the board (Sackett et al., 2022). And unlike most top predictors, the upgrade from weak to strong costs process rather than money — the same interviewer hours and the same candidate hours, reorganized around job-derived questions and anchored scoring. Few findings in applied psychology offer this much validity per unit of disruption. That makes the unstructured interview’s persistence not just a scientific puzzle but a standing free lunch, declined daily.
Under the most conservative modern corrections, the structured interview emerged as the single strongest common predictor of job performance — and unlike most top predictors, the upgrade costs process rather than money: the same hours, reorganized around job-derived questions and anchored scoring (Sackett et al., 2022).
The belief that seasoned intuition can out-predict a structured protocol is selection’s most persistent — and most expensive — myth.
Paraphrasing Highhouse (2008), Industrial and Organizational Psychology
Where BARS came from
Half of structure is which questions get asked. The other half is what happens to the answers. That half has its own history — one that began not in hiring at all but in the measurement of nursing, with a build method so careful it still embarrasses most modern rubrics. The origin is Smith and Kendall’s paper on the “retranslation of expectations,” written to repair rating scales for nursing performance (Smith & Kendall, 1963). Their diagnosis: scale points labeled with adjectives — outstanding, average, poor — mean different things to different raters. So a rating mixes the ratee’s behavior with the rater’s private vocabulary.
Their build method is the interesting part. Experts on the job write critical incidents — short accounts of effective and ineffective behavior — for each performance dimension. A second, independent panel then retranslates each incident: it assigns the incident back to the dimension it supposedly illustrates. Any incident the panel cannot reliably reassign is thrown out. The survivors are scaled for effectiveness, and only incidents that judges place at the same level with high agreement become anchors. What remains is a behaviorally anchored rating scale — a dimension whose score points are defined by concrete, observable behaviors that independent judges agree show that dimension at that level.
The point of all this machinery is shared meaning. “Communicates well” invites projection; a description of a specific, observable action does not. Retranslation forces raters into a common frame of reference before the first candidate walks in. That is why BARS became the standard scoring half of the structured interview.
Why structure wins
A validity gap this large and this stable demands a mechanism, not just a number. Partly for scientific hygiene; partly because knowing why structure wins tells a program which parts it cannot afford to drop. Three mechanisms explain the gap. None of them requires believing that interviewers are fools.
First, reliability. A predictor’s validity is capped, as a matter of math, by its reliability — how consistently it scores the same thing. An interview that two raters score differently cannot predict anything well. A meta-analysis of interview reliability, both between raters and within the interview, found that rater agreement rises sharply with structure (Conway, Jako & Goodman, 1995). Standardized questions and anchored, per-question ratings produce far more consistent scores. That lifts the ceiling on what the interview can predict.
Second, comparability. When every candidate answers the same job-derived questions, differences in ratings reflect differences in answers. When each interview wanders its own path, no two candidates produce evidence that can be compared. The rating then becomes a contest of impressions rather than of performances — friendly ground for confirmation bias, halo, and similar-to-me effects.
Third, decomposition. Rating each answer against an anchor, then combining scores by rule, prevents one vivid moment — good or bad — from colonizing the whole evaluation. Anchors change the rater’s task from magnitude estimation (how good did that feel?) to matching (which described behavior does this answer most resemble?), and matching is a task human judges perform far more consistently.
The standard objection — that structure kills rapport and turns interviews into interrogations — misreads what the parts require. Nothing in the structure literature forbids warmth, welcome, or a genuinely human conversation around the assessment core. What it constrains is which questions carry evidential weight and how the answers are scored. The rapport-building minutes can stay; they just stop being the measurement. And candidates report, again and again, that well-run structured interviews feel fair — everyone faced the same questions, judged the same way (Levashina et al., 2014). That is exactly the claim the free-form chat cannot honestly make.
The defensibility case
For companies unmoved by validity numbers, the literature offers a second currency. Validity is half the case for structure; the other half is legal. One study analyzed federal employment-discrimination cases that involved interviews. Verdicts tended to favor the defending company when the interview had structure: objective, job-related criteria, the same process for everyone, and written ratings a court can review (Williamson et al., 1997). An anchored score trail is evidence of a job-related process. A gut call is an assertion.
The fairness evidence points the same direction. The Levashina review collects findings that structured interviews tend to show smaller subgroup differences than many rival predictors, while staying among the most valid (Levashina et al., 2014). That pairing is unusual in selection, where validity and adverse impact so often trade off. The mechanism is the same one that lifts reliability. Structure shrinks the discretionary channel through which similarity, comfort, and stereotype travel. What remains in the score is more answer and less rater.
Keeping structure from decaying
Structure is not a one-time install; it is upkeep, because every force in the room pushes back toward chat. The decay pattern is well known to anyone who has audited a mature program. Interviewers begin to improvise follow-ups that leak the anchors. Favorite questions drift in and the job-analysis questions drift out. Panels start “discussing to consensus” — which restores exactly the impression-forming and rank effects that per-rater, rule-combined scoring exists to prevent (Campion, Palmer & Campion, 1997). Within a couple of hiring cycles, a company can be running Level 4 paperwork around Level 2 practice, enjoying the validity of neither.
Structure decays. Improvised follow-ups leak the anchors, favorite questions displace the job-analysis ones, and “discussing to consensus” restores the impression-forming the protocol was built to prevent — the paperwork stays Level 4 while the practice slides (Campion, Palmer & Campion, 1997).
The fixes come straight from the components. Refresh rater training on the anchors — the frame-of-reference habit that decays and retrains well. Audit the score data, not the forms. Ratings that move in lockstep within one rater signal halo creeping back; raters whose candidates cluster at one scale point signal private standards returning. Keep the combination mechanical, with any override written down — the rule the combination research supplies (Conway, Jako & Goodman, 1995). And treat every new interviewer’s first cycles as calibration: score them in parallel against a seasoned anchor-rater before their ratings count alone.
None of this is exotic. It is the same instrument upkeep any measurement system needs — applied to a measurement system that happens to be made of people.
How Future Proof™ applies this.
Every competency in a Future Proof interview carries a behaviorally anchored rubric: concrete descriptions of what a weak, adequate, and strong answer actually sounds like, written per competency and per level. Answers are scored by matching evidence to anchors — never by overall impression. Each competency has a defined threshold, and critical competencies carry gates: a below-gate score cannot be offset by charisma elsewhere in the interview. It is the interview the validity literature describes — same questions, anchored judgment, decomposed scoring — run consistently at scale.
See the interview module →What the evidence doesn’t show
The literature is strong, but it is routinely oversold in four ways:
- BARS is not a magic format. Reviewing two decades of rating-format research, Landy and Farr concluded that format changes alone — including BARS versus other carefully built scales — yield modest psychometric gains, and famously proposed a moratorium on format research in favor of studying raters and rating processes (Landy & Farr, 1980). Anchors earn their keep as part of a system — job analysis, retranslation, rater training, decomposed scoring — not as a template pasted over old habits.
- More structure is not monotonically better. Validity plateaus at the top of the structure scale (Huffcutt & Arthur, 1994): once questions are standardized and scoring is anchored, additional rigidity buys little.
- The exact numbers are estimates, not constants. The 2022 re-analysis showed that widely quoted validity figures had been inflated by aggressive range-restriction corrections (Sackett et al., 2022). Treat any single coefficient as a point estimate with real uncertainty; it is the ordering of methods that has proven robust.
- Structure does not eliminate impression management. The Levashina review documents that applicants fake and self-promote in structured interviews too; question format constrains it without removing it (Levashina et al., 2014). And structure decays in the field — interviewers drift back toward conversation unless the process holds the line.
Where the evidence stops
- 1BARS is not a magic format
- 2More structure is not monotonically better
- 3The exact numbers are estimates, not constants
- 4Structure does not eliminate impression management
None of these caveats disturbs the central result. The structured interview with anchored scoring is that rare tool that is at once more valid, more reliable, fairer, and more defensible than the intuitive alternative. It has been all four for as long as anyone has measured.
Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The interview and rating literatures are among the deepest in applied psychology; these twelve are load-bearing.
The evidence, by year
- 1963Smith
- 1980Landy
- 1988Wiesner
- 1994Huffcutt
- 1994McDaniel
- 1995Conway
- 1997Campion
- 1997Williamson
- 1998Schmidt
- 2008Highhouse
- 2014Levashina
- 2022Sackett
- Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology 1(3): 333–342. PDF
- Levashina, J., Hartwell, C.J., Morgeson, F.P., & Campion, M.A. (2014). The structured employment interview: Narrative and quantitative review of the research literature. Personnel Psychology 67(1): 241–293. DOI
- Campion, M.A., Palmer, D.K., & Campion, J.E. (1997). A review of structure in the selection interview. Personnel Psychology 50(3): 655–702. PDF
- Huffcutt, A.I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology 79(2): 184–190. DOI
- Wiesner, W.H., & Cronshaw, S.F. (1988). A meta-analytic investigation of the impact of interview format and degree of structure on the validity of the employment interview. Journal of Occupational Psychology 61(4): 275–290. PDF
- McDaniel, M.A., Whetzel, D.L., Schmidt, F.L., & Maurer, S.D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology 79(4): 599–616. DOI
- Schmidt, F.L., & Hunter, J.E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin 124(2): 262–274. DOI
- Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOI
- Smith, P.C., & Kendall, L.M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology 47(2): 149–155. DOI
- Landy, F.J., & Farr, J.L. (1980). Performance rating. Psychological Bulletin 87(1): 72–107. DOI
- Conway, J.M., Jako, R.A., & Goodman, D.F. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied Psychology 80(5): 565–579. DOI
- Williamson, L.G., Campion, J.E., Malos, S.B., Roehling, M.V., & Campion, M.A. (1997). Employment interview on trial: Linking interview structure with litigation outcomes. Journal of Applied Psychology 82(6): 900–912. DOI
Interviews scored against anchors, not impressions.
Book a 20-minute demo with one of your real roles. We’ll walk through a competency rubric, the anchors behind each score, and exactly where the thresholds and gates sit — so every hiring decision arrives with its evidence attached.