© 2026 FUTURE PROOF™
Assessment Science · Feedback

Feedback: the intervention that can backfire.

Everyone believes in feedback. The most famous meta-analysis in the field found that more than a third of feedback interventions made performance worse. What the evidence says about the difference — and the feedback rules Future Proof™’s engine is built on.

TL;DR

The finding: Feedback improves learning on average (meta-analytic d ≈ 0.4–0.5) — and the average hides enormous variance. In Kluger & DeNisi’s landmark synthesis of 607 effect sizes, over one third were negative. The strongest moderator is where the feedback points the learner’s attention: at the task and how to improve it, or at the self.

The mechanism: Feedback that carries information — what was wrong, why, what to do next — teaches. Feedback that carries judgment — praise, person-level verdicts, bare scores — pulls the mind to the self and crowds out the task. Elaborated feedback beats correct-answer feedback, which beats right/wrong marks. And all of it only works if the learner attempts first and processes the correction.

The product: Future Proof’s engine delivers task-level, elaborated feedback — the why and the next step, never person-level judgment — after a genuine attempt, then verifies the correction stuck by re-asking later. Scores exist for dashboards; explanations exist for learners.

In this article

  1. 01Where the feedback points
  2. 02The four levels, on one wrong answer
  3. 03Information beats verdicts
  4. 04Feedback is received, not delivered
  5. 05The timing question
  6. 06Why platforms defaulted to the weakest kind — and why that’s over
  7. 07The workplace corollary: why performance reviews keep failing
  8. 08What the evidence doesn’t show
  9. 09What this means for practice
© 2026 FUTURE PROOF™
The route. 9 sections, from “Where the feedback points” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Every learning system, every management framework, and every parenting book agrees on one tool. Pause on that before we examine it. Agreement this complete is rare in the behavioral sciences. It should also make us wary. The tool’s flagship meta-analysis — a study that pools the results of many experiments — is best known for an awkward finding: the tool often backfires.

No teaching belief is safer to say out loud than “people need feedback.” Managers, teachers, and platform designers all agree on it. That is what makes the central paper of the feedback literature so uncomfortable. Kluger and DeNisi gathered nearly a century of feedback experiments: 607 effect sizes across thousands of people. (An effect size is a standard measure of how big an impact is.) The average effect was positive but modest (d ≈ 0.41) — and more than one in three interventions made performance worse (Kluger & DeNisi, 1996).

The number

> 1 in 3 The share of the 607 feedback interventions that made performance worse — the most agreed-upon tool in learning and management, routinely damaging the thing it exists to improve (Kluger & DeNisi, 1996).

Not “failed to help.” Worse. The most agreed-upon tool in learning and management routinely damages the very thing it is meant to improve. The rest of the literature is, in effect, one long hunt for which third that is. The answer turns out to be one of the most useful findings in applied psychology.

Where the feedback points

Kluger and DeNisi’s explanation is called feedback intervention theory, and every later synthesis upholds it. The effect of feedback depends on where it points the learner’s attention. Feedback aimed at the task — what the correct response was, what process produces it — helps. Feedback that pulls attention to the self — praise, criticism, anything read as a verdict on the person — reliably shrinks or reverses the benefit (Kluger & DeNisi, 1996). Why? The learner’s mental resources go to managing the judgment instead of closing the gap.

The framework most educators know sorts this into four levels: feedback about the task, about the process, about self-regulation, and about the self (Hattie & Timperley, 2007). The first three carry nearly all the value. The fourth is the most common in practice — and it carries nearly none.

The praise studies make the self-level danger concrete. Children praised for intelligence after a success later chose easier problems, enjoyed them less, and collapsed harder after failure than children praised for effort (Mueller & Dweck, 1998). Even flattering self-level feedback is a loan against future difficulty.

The four levels, on one wrong answer

The framework turns vivid the moment you apply it to one concrete miss. Say a learner computes a project’s break-even point and gets it wrong. Task-level feedback says: the answer is 4,200 units, and you divided fixed costs by price instead of by contribution margin. Process-level feedback says: when a break-even answer looks too low, check whether variable costs made it into the denominator. That check carries over to every problem of this family.

Self-regulation-level feedback says: you answered in eleven seconds with high confidence. On multi-step calculations, your fast-and-sure answers are wrong twice as often as your slow ones — so consider checking before you commit. Self-level feedback says: great effort, you’re nearly there!

Read in order, the levels form a ladder: each level transfers further than the last. The task level fixes this answer. The process level fixes this family of problems. The self-regulation level fixes the learner’s relationship to their own accuracy — the calibration loop our metacognition review covers. The fourth level fixes nothing, yet most corporate tools default to it, because it is the only one that can be written without knowing anything about the task. That is the tell worth keeping: feedback that could be pasted under any answer is feedback about nothing.

no effect ≈ 38% of 607 effects negative majority positive · mean d ≈ 0.41 Feedback intervention effects, Kluger & DeNisi (1996) © 2026 FUTURE PROOF™
Figure 1. The distribution behind the average. Feedback helps on average — and over a third of the 607 effect sizes in the landmark meta-analysis were negative. The moderator analysis, not the mean, is the finding. Schematic; proportions after Kluger & DeNisi (1996). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Information beats verdicts

Hold attention on the task, and the next question is what to say. Here the research on computer-based lessons is unusually clean, because feedback content can be varied with precision. The ranking repeats across meta-analyses. Elaborated feedback explains why the answer is right or wrong, ideally with a pointer to the next step. It beats correct-answer feedback, which in turn beats bare right/wrong marks — and the edge of elaboration is largest for higher-order outcomes (Van der Kleij, Feskens & Eggen, 2015).

The most recent large synthesis puts numbers on it. Feedback overall averages d ≈ 0.48 — a moderate effect. High-information feedback approaches d ≈ 0.99, roughly the difference between a weak intervention and one of the strongest known (Wisniewski, Zierer & Hattie, 2020).

A classroom study from the grades literature makes the same point through a side door. Students given comments alone improved. Students given grades — or even grades plus comments — did not (Butler, 1988). The verdict swallowed the information beside it. A score is where a learner stops reading.

Notice what that result does not say. It does not say scores are useless. It says they are poison at the moment of learning. Managers need summary numbers; learners mid-lesson need explanations. The mistake is serving both audiences the same screen.

The design translation is a split of duties. Percentages, grades, and leaderboards live in dashboards, checked between sessions. The in-lesson channel carries nothing but the information the next attempt can use. Platforms that stamp a running score on every question have, per Butler’s data, installed a device that teaches learners to stop reading the part that teaches.

Design rule

Separate the audiences. Scores, grades, and leaderboards belong in dashboards consulted between sessions; the in-lesson channel carries only what the next attempt can use. A verdict placed beside an explanation swallows it — so never make a learner read past a number to reach the information.

Feedback is received, not delivered

An older synthesis added a condition that modern platforms still break daily. Feedback only works when the learner mindfully processes it. And it is reliably neutralized by “pre-search availability” — letting the learner see the answer before truly attempting the question (Bangert-Drowns, Kulik, Kulik & Morgan, 1991). Feedback presupposes a committed response to correct. A learner who can peek has nothing at stake, and learns accordingly. The same line of reviews established that feedback is not behaviorist reinforcement — it is information the learner must do something with (Kulhavy, 1977).

The receiving side has its own design surface, and most of it is about engineering a moment of real processing. Commitment comes first: no answer, no reveal — the interface rule that puts Bangert-Drowns into practice. Then put friction in the right place. A “got it” button costs nothing and certifies nothing. Asking the learner to apply the correction — retry the item, answer a near-transfer variant, or restate the rule in their own words — forces the mindful processing the meta-analysis found essential.

The learner’s confidence at the moment of answering is worth capturing too. A confident wrong answer marks a structured misconception rather than a blank. So it deserves a fuller explanation than a tentative one. The hypercorrection literature says that is exactly where a good correction pays most. None of this is expensive. All of it is the difference between feedback as a message sent and feedback as a message received.

The timing question

Timing turns out to be the least dramatic variable, though the seeming clash in the literature confused practitioners for years. In real classroom settings, immediate feedback tends to beat delayed (Kulik & Kulik, 1988). In lab studies of retention, a delay can help — in effect, it gives the learner a second spaced pass at the material (Butler, Karpicke & Roediger, 2007).

The two findings stop conflicting once you notice what a delay costs in each setting. In the lab, the delayed feedback is guaranteed to be read. In a classroom or a compliance course, feedback that arrives next week reaches a learner who no longer remembers the question, no longer cares, and often never opens it. Delay’s memory benefit is real but fragile. Its engagement cost in the field is large and certain.

The practical fix captures both. Deliver the explanation while the question still matters to the learner. Then re-test the corrected item days later. That turns the correction into retrieval practice. It supplies the spaced second exposure the delay studies were really measuring, and it produces evidence that the fix held. Immediate explanation plus a scheduled re-test is not a compromise between the two literatures — it is the design both were pointing at.

Measured, in Cohen’s d High-information All feedback Grand mean, 19960.99 0.48 0.410 0.25 0.5 0.75 1.0 Effect size (d) Rank order — not to scale Right/wrong marks Correct answer only Elaborated feedback Information content, ranked © 2026 FUTURE PROOF™
Figure 2. Information content is the moderator. Left, measured averages: feedback overall runs d ≈ 0.48 against d ≈ 0.99 for high-information feedback (Wisniewski, Zierer & Hattie, 2020), with the century-long grand mean of d ≈ 0.41 (Kluger & DeNisi, 1996) plotted for scale. Right, the content ranking that repeats across meta-analyses: elaborated feedback beats correct-answer feedback, which beats bare right/wrong marks (Van der Kleij, Feskens & Eggen, 2015). Step lengths in the right-hand panel are ordinal, not measured. Meta-analytic averages simplify wide distributions — read the contrast, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Why platforms defaulted to the weakest kind — and why that’s over

If elaborated feedback wins so clearly, why has courseware served right/wrong marks for forty years? Economics, not ignorance. A red X costs nothing to generate. An explanation of this learner’s particular error once required a human who understood both the content and the mistake — affordable in a tutoring session, impossible across ten thousand quiz attempts. Correct-answer feedback was the compromise the technology could afford. So the meta-analytic finding that it underperforms elaboration (Van der Kleij et al., 2015) was, for practical purposes, a finding about a price nobody could pay.

That constraint has now collapsed. Writing a specific, task-level explanation for a specific wrong answer is exactly the kind of work modern AI systems do well, at marginal costs near zero. That moves the design question from “can we afford elaboration?” to “is the elaboration any good?” And the quality bar the literature sets is concrete.

The explanation must engage the learner’s actual error, not recite the textbook paragraph the item came from. It must stay at task and process level. And it must be right. A confidently wrong explanation, delivered at the moment of maximum attention, is the one failure mode worse than a bare X. Generated feedback therefore needs the same validation discipline as generated questions. It must be grounded in the item’s rationale, spot-audited by humans, and watched through the retest data that shows whether corrections are actually landing.

The catch

Cheap elaboration raises the stakes of being wrong. A confidently incorrect explanation, delivered at the moment of maximum attention, is the one failure mode worse than a bare X — so generated feedback inherits the full validation discipline of generated questions, not a lighter one.

One of the most powerful influences on learning — and one of the most variable. The double-edged conclusion of the feedback meta-analyses, after Hattie & Timperley (2007).

The workplace corollary: why performance reviews keep failing

It is worth remembering where Kluger and DeNisi’s corpus came from. Not classrooms — a century of workplace and laboratory performance interventions. The annual review is a delayed, pooled, person-level verdict, delivered with a rating attached. It manages to sit at the wrong end of nearly every moderator — every helps-or-hurts factor — in their table at once.

Count the ways. The information arrives months after the behavior it describes. It is pooled past the point where any specific action could be pulled from it. It is framed as a judgment of the person. And it is anchored to a score — which, per Butler’s classroom result, is where the reading stops. The genre’s persistent failure to improve performance is not a mystery the feedback literature cannot explain; it is the literature’s central prediction, running every year in most large organizations.

The repair path is the same one the learning side prescribes, translated: shrink the loop, and drop the verdict a level. Tie feedback to a specific recent piece of work. Describe what happened and what to do differently next time, close to the event — task and process level, in the flow. That is the workplace shape of elaborated feedback, and it is precisely what the better-evidenced coaching practices consist of. The annual aggregate can survive as a pay instrument if it must. But expecting it to also be the improvement engine is asking a verdict to do information’s job — the exact confusion this literature spent a hundred years documenting.

What the evidence doesn’t show

  • Feedback is not a substitute for instruction. Formative-feedback guidelines are explicit that feedback works on top of a genuine attempt at a learnable task; it cannot rescue content the learner never engaged (Shute, 2008).
  • More is not better. Continuous correction during early skill acquisition can create dependency, with learners performing well under feedback and collapsing without it — one reason formative guidelines recommend against interrupting a learner mid-attempt (Shute, 2008).
  • The meta-analytic base skews short-term and academic. Most effects are measured on near-term performance in instructional settings; the workplace performance-review literature is far messier, and Kluger & DeNisi’s warnings apply there with extra force.
  • Praise is not banned — it is just not feedback. The evidence objects to praise as the information channel. Warmth and encouragement matter for persistence; they simply cannot replace the task-level content that does the teaching.

Where the evidence stops

  1. 1Feedback is not a substitute for instruction
  2. 2More is not better
  3. 3The meta-analytic base skews short-term and academic
  4. 4Praise is not banned — it is just not feedback
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

Audit your feedback the way you would audit content, because it is content. Arguably it is the highest-leverage content in the system — it lands at the one moment the learner is sure to be paying attention. Walk through a real session. Sort every piece of feedback a learner meets into the four levels. Most audits find the same spread: a thin layer of task-level correction, almost no process or self-regulation content, and a thick blanket of person-level cheerleading and scores. Per thirty years of meta-analysis, that spread is close to the least effective setup available.

The rebuild follows directly from the evidence. Require a committed attempt before any answer is visible. Say what was wrong, why, and what to try next — task and process level, never the person. Prefer explanations to scores everywhere a learner is still learning. Quarantine the scores in dashboards, where they can inform without interrupting.

Deliver the explanation while the question is warm. Then close the loop the way the retention literature demands: schedule the corrected item to return. A correction that is never re-tested is an assumption, not an outcome. And treat “we give learners feedback” as the start of the design conversation, not the end. The meta-analyses are unanimous: the sentence means nothing until you specify what kind, at what level, verified how.

Applied research

How Future Proof™ applies this: feedback with information in it.

Every question in the engine requires a committed attempt before anything is revealed — no pre-search, no peeking. What follows is elaborated, task-level feedback: why the chosen answer fails, what distinguishes the correct one, and what to look at next, generated against the specific error the learner made. No person-level judgment, no bare scores mid-lesson. Then the engine closes the loop the literature says to close: corrected items re-enter the spaced schedule, so the fix is verified by a later retrieval, not assumed from a click on “got it.”

See the AI Engine
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 1977Kulhavy
  • 1988Butler
  • 1988Kulik
  • 1991Bangert-Drowns
  • 1996Kluger
  • 1998Mueller
  • 2007Hattie
  • 2007Butler
  • 2008Shute
  • 2015Kleij
  • 2020Wisniewski
© 2026 FUTURE PROOF™
The evidence base. The 11 sources cited here span 1977–2020, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Kluger, A.N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin 119(2): 254–284. DOI
  2. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research 77(1): 81–112. DOI
  3. Mueller, C.M., & Dweck, C.S. (1998). Praise for intelligence can undermine children’s motivation and performance. Journal of Personality and Social Psychology 75(1): 33–52. DOI
  4. Van der Kleij, F.M., Feskens, R.C.W., & Eggen, T.J.H.M. (2015). Effects of feedback in a computer-based learning environment on students’ learning outcomes: A meta-analysis. Review of Educational Research 85(4): 475–511. PDF
  5. Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology 10: 3087. DOI
  6. Butler, R. (1988). Enhancing and undermining intrinsic motivation: The effects of task-involving and ego-involving evaluation on interest and performance. British Journal of Educational Psychology 58(1): 1–14. PDF
  7. Bangert-Drowns, R.L., Kulik, C.-L.C., Kulik, J.A., & Morgan, M. (1991). The instructional effect of feedback in test-like events. Review of Educational Research 61(2): 213–238. PDF
  8. Kulhavy, R.W. (1977). Feedback in written instruction. Review of Educational Research 47(2): 211–232. PDF
  9. Kulik, J.A., & Kulik, C.-L.C. (1988). Timing of feedback and verbal learning. Review of Educational Research 58(1): 79–97. PDF
  10. Butler, A.C., Karpicke, J.D., & Roediger, H.L. (2007). The effect of type and timing of feedback on learning from multiple-choice tests. Journal of Experimental Psychology: Applied 13(4): 273–281. PDF
  11. Shute, V.J. (2008). Focus on formative feedback. Review of Educational Research 78(1): 153–189. DOI
Try the AI engine

See feedback with information in it.

Book a 20-minute demo with your team’s actual content. We’ll show you the elaborated feedback the engine generates for real wrong answers — and the retest that proves the correction stuck.

11 citations Reviewed August 2026 Open peer review welcomed