Research · Assessment Science
Assessment Science · Structured Interviews

Structured interviews and BARS: the validity case.

Hiring managers trust their read of a candidate. Sixty years of selection research say the read is the problem. On the evidence for structured interviews and behaviorally anchored rating scales — the same architecture Future Proof™’s interview module scores against, with anchors, thresholds, and gates instead of gut feel.

TL;DR

The finding: Structured interviews carry roughly double the predictive validity of unstructured ones. The result has held from the first meta-analyses in the late 1980s through the most conservative re-analysis in 2022 — where the structured interview ranked as the single best predictor of job performance among common selection tools.

The mechanism: Structure removes noise from judgment. Every candidate gets the same job-derived questions; every answer is rated against behaviorally anchored scales; scores combine by rule. Anchors turn evaluation from impression into evidence-matching, and standardization makes candidates comparable — both raise reliability, which caps validity.

The product: Future Proof’s interview module scores every answer against per-competency BARS anchors with defined thresholds and gates — not gut impressions.

The employment interview is the most widely used selection method in the world, and the version most organizations use — an open conversation steered by the interviewer’s instincts — is the version the evidence supports least. That is not a new finding, and it is not a close call. Across six decades of meta-analytic work the pattern is stable: interviews predict job performance well when they are structured and poorly when they are not, and the gap between the two is one of the largest method effects in personnel psychology.

The persistence of the unstructured interview is itself a research topic. Highhouse documented what he called practitioners’ stubborn reliance on intuition and subjectivity in selection — the belief that an experienced judge, freed from constraints, can read a candidate better than any protocol (Highhouse, 2008). Decision research points the other way, and so, very specifically, does the interview literature.

What “structure” actually means

Structure is not a synonym for rigid or scripted. In the definitive modern review, Levashina, Hartwell, Morgeson, and Campion treat structure as any enhancement of the interview intended to improve its measurement properties by standardizing what is asked, how answers are evaluated, or both (Levashina et al., 2014).

An earlier review catalogued fifteen distinct components of structure and sorted them into two families (Campion, Palmer & Campion, 1997). Content structure governs the questions: base them on a job analysis, ask every candidate the same ones, prefer behavioral and situational formats, limit ad-hoc prompting. Evaluation structure governs the judging: rate each answer rather than forming a single overall impression, anchor the rating scale in described behaviors, take notes, use multiple interviewers, and combine scores by formula rather than by discussion.

Because structure is a bundle of components, it is a continuum rather than a switch. Huffcutt and Arthur classified interview studies into four ascending levels of structure, from no constraints on questions or scoring to fully standardized questions scored on anchored scales (Huffcutt & Arthur, 1994). That framing turned “do interviews work?” into a better question: how does validity move as structure increases?

.60 .45 .30 .15 .20 .35 .56 .57 plateau Level 1 Level 2 Level 3 Level 4 no constraints limited high complete Validity
Figure 1. Mean corrected interview validity at four levels of structure. Validity nearly triples from the least to the most structured formats, then plateaus at the top. Values from Huffcutt & Arthur (1994).

The validity record

The answer arrived in waves of meta-analysis. Wiesner and Cronshaw found that structured interviews carried roughly twice the predictive validity of unstructured ones — corrected coefficients of about .63 versus .20 (Wiesner & Cronshaw, 1988). Working from a larger database with more conservative assumptions, McDaniel, Whetzel, Schmidt, and Maurer found the same ordering at lower magnitudes: about .44 for structured interviews against job-performance criteria, versus about .33 for unstructured (McDaniel et al., 1994). The level analysis filled in the shape of the curve: validity climbed from roughly .20 at the lowest level of structure to roughly .57 at the highest, with most of the gain arriving by the second-highest level (Huffcutt & Arthur, 1994).

Summarizing eighty-five years of selection research, Schmidt and Hunter put the structured interview at .51 — level with general mental ability among the strongest single predictors available — against .38 for the unstructured version (Schmidt & Hunter, 1998). And when Sackett, Zhang, Berry, and Lievens re-estimated the whole table with more defensible corrections for range restriction, nearly every predictor’s validity fell, but the ordering sharpened: the structured interview emerged at the top of the revised rankings, at roughly .42, while the unstructured interview dropped far down the list (Sackett et al., 2022).

The exact coefficients move with the correction method. The direction never has.

The belief that seasoned intuition can out-predict a structured protocol is selection’s most persistent — and most expensive — myth.

Paraphrasing Highhouse (2008), Industrial and Organizational Psychology

Where BARS came from

The evaluation half of structure has a precise origin: Smith and Kendall’s paper on the “retranslation of expectations,” written to repair rating scales for nursing performance (Smith & Kendall, 1963). Their diagnosis was that scale points labeled with adjectives — outstanding, average, poor — mean different things to different raters, so a rating mixes the ratee’s behavior with the rater’s private vocabulary.

Their construction method is the interesting part. Subject-matter experts generate critical incidents of effective and ineffective behavior on each performance dimension. A second, independent panel then retranslates each incident — assigns it back to the dimension it supposedly illustrates — and any incident the panel cannot reliably reassign is discarded. The survivors are scaled for effectiveness, and only incidents that judges place at the same level with high agreement become anchors. What remains is a behaviorally anchored rating scale: a dimension whose score points are defined by concrete, observable behaviors that independent judges agree exemplify that dimension at that level.

The point of all this machinery is shared meaning. “Communicates well” invites projection; a description of a specific, observable action does not. Retranslation forces raters into a common frame of reference before the first candidate walks in — which is why BARS became the canonical evaluation component of the structured interview.

Why structure wins

Three mechanisms explain the validity gap, and none of them requires believing that interviewers are fools.

First, reliability. A predictor’s validity is mathematically capped by its reliability: an interview that two raters score differently cannot predict anything well. A meta-analysis of interrater and internal-consistency reliability across interview studies found that agreement between raters rises substantially with structure — standardized questions and anchored, per-question ratings produce far more consistent scores, lifting the ceiling on what the interview can predict (Conway, Jako & Goodman, 1995).

Second, comparability. When every candidate answers the same job-derived questions, differences in ratings reflect differences in answers. When each interview wanders its own path, candidates generate incommensurable evidence, and the rating becomes a comparison of impressions rather than of performances — hospitable terrain for confirmation bias, halo, and similar-to-me effects.

Third, decomposition. Rating each answer against an anchor, then combining scores by rule, prevents one vivid moment — good or bad — from colonizing the whole evaluation. Anchors change the rater’s task from magnitude estimation (how good did that feel?) to matching (which described behavior does this answer most resemble?), and matching is a task human judges perform far more consistently.

The defensibility case

Validity is half the argument for structure; the other half is legal. An analysis of federal employment-discrimination cases involving interviews found that structural characteristics — objective, job-related criteria; standardized administration; documented, reviewable evaluations — were associated with verdicts favoring the defending organization (Williamson et al., 1997). An anchored score trail is evidence of a job-related process; a gut call is an assertion.

The fairness evidence points the same direction. The Levashina review collects findings that structured interviews tend to show smaller subgroup differences than many alternative predictors while remaining among the most valid — an unusual combination in selection, where validity and adverse impact so often trade off (Levashina et al., 2014).

Applied in practice

How Future Proof™ applies this.

Every competency in a Future Proof interview carries a behaviorally anchored rubric: concrete descriptions of what a weak, adequate, and strong answer actually sounds like, written per competency and per level. Answers are scored by matching evidence to anchors — never by overall impression. Each competency has a defined threshold, and critical competencies carry gates: a below-gate score cannot be offset by charisma elsewhere in the interview. It is the interview the validity literature describes — same questions, anchored judgment, decomposed scoring — run consistently at scale.

See the interview module

What the evidence doesn’t show

The literature is strong, but it is routinely oversold in four ways:

  • BARS is not a magic format. Reviewing two decades of rating-format research, Landy and Farr concluded that format changes alone — including BARS versus other carefully built scales — yield modest psychometric gains, and famously proposed a moratorium on format research in favor of studying raters and rating processes (Landy & Farr, 1980). Anchors earn their keep as part of a system — job analysis, retranslation, rater training, decomposed scoring — not as a template pasted over old habits.
  • More structure is not monotonically better. Validity plateaus at the top of the structure scale (Huffcutt & Arthur, 1994): once questions are standardized and scoring is anchored, additional rigidity buys little.
  • The exact numbers are estimates, not constants. The 2022 re-analysis showed that widely quoted validity figures had been inflated by aggressive range-restriction corrections (Sackett et al., 2022). Treat any single coefficient as a point estimate with real uncertainty; it is the ordering of methods that has proven robust.
  • Structure does not eliminate impression management. The Levashina review documents that applicants fake and self-promote in structured interviews too; question format constrains it without removing it (Levashina et al., 2014). And structure decays in the field — interviewers drift back toward conversation unless the process holds the line.

None of these caveats disturbs the central result. The structured interview with anchored scoring is that rare instrument that is simultaneously more valid, more reliable, fairer, and more defensible than the intuitive alternative — and it has been all four for as long as anyone has measured.

References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The interview and rating literatures are among the deepest in applied psychology; these twelve are load-bearing.

  1. Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology 1(3): 333–342. PDF
  2. Levashina, J., Hartwell, C.J., Morgeson, F.P., & Campion, M.A. (2014). The structured employment interview: Narrative and quantitative review of the research literature. Personnel Psychology 67(1): 241–293. DOI
  3. Campion, M.A., Palmer, D.K., & Campion, J.E. (1997). A review of structure in the selection interview. Personnel Psychology 50(3): 655–702. PDF
  4. Huffcutt, A.I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology 79(2): 184–190. DOI
  5. Wiesner, W.H., & Cronshaw, S.F. (1988). A meta-analytic investigation of the impact of interview format and degree of structure on the validity of the employment interview. Journal of Occupational Psychology 61(4): 275–290. PDF
  6. McDaniel, M.A., Whetzel, D.L., Schmidt, F.L., & Maurer, S.D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology 79(4): 599–616. DOI
  7. Schmidt, F.L., & Hunter, J.E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin 124(2): 262–274. DOI
  8. Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology 107(11): 2040–2068. DOI
  9. Smith, P.C., & Kendall, L.M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology 47(2): 149–155. DOI
  10. Landy, F.J., & Farr, J.L. (1980). Performance rating. Psychological Bulletin 87(1): 72–107. DOI
  11. Conway, J.M., Jako, R.A., & Goodman, D.F. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied Psychology 80(5): 565–579. DOI
  12. Williamson, L.G., Campion, J.E., Malos, S.B., Roehling, M.V., & Campion, M.A. (1997). Employment interview on trial: Linking interview structure with litigation outcomes. Journal of Applied Psychology 82(6): 900–912. DOI
See it in practice

Interviews scored against anchors, not impressions.

Book a 20-minute demo with one of your real roles. We’ll walk through a competency rubric, the anchors behind each score, and exactly where the thresholds and gates sit — so every hiring decision arrives with its evidence attached.

12 citations Reviewed July 2026 Open peer review welcomed