© 2026 FUTURE PROOF™
Memory & Practice · Pretesting

The pretesting effect: wrong answers that teach.

Asking learners questions they cannot yet answer looks like a waste of everyone’s time. The evidence says the failed attempt primes the memory system for what comes next — and that errorless instruction is the real waste. Why Future Proof™ opens every module with questions, not content.

TL;DR

The finding: Attempting to answer a question before the material has been taught — and almost always failing — produces better learning of the subsequently studied answer than spending the same time studying. The result holds for word pairs, prose passages, and video lectures, and it survives systematic review.

The mechanism: A retrieval attempt activates whatever related knowledge the learner has, exposes the gap where the answer should be, and sharpens attention to the answer when it arrives. Errors help rather than hurt — provided corrective feedback follows; high-confidence errors are corrected best of all.

The product: Future Proof’s modules are diagnostic-first: the adaptive engine asks before it teaches, uses the errors to map what each learner is missing, and delivers the content as the answer to questions the learner has just cared about getting wrong.

In this article

  1. 01The errorless inheritance
  2. 02Guessing wrong on purpose
  3. 03Why failure primes learning
  4. 04What counts as an attempt
  5. 05The pretest as an instrument, not just an intervention
  6. 06Pretesting vs. posttesting
  7. 07The feedback condition
  8. 08What the evidence doesn’t show
  9. 09What this means for practice
© 2026 FUTURE PROOF™
The route. 9 sections, from “The errorless inheritance” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Every course has an opening move, and nearly all of them choose the same one: explain. The syllabus overview, the concept video, the walkthrough. Instruction first, questions later — on the theory that you cannot fairly ask people what they have not been taught. This article is about the experiments that tested the opposite move. They kept winning with it.

Teaching common sense says errors are damage. You explain first, carefully, so that learners never practice a mistake. On this view, a question asked before the lesson is at best a gimmick — at worst, a way to rehearse wrong answers. Behaviorist theory made this a principle: errorless learning. Most corporate training still follows it — content first, quiz after.

Over the past two decades, a line of experiments has flipped that sequence and measured what happens. Learners who attempt answers before studying — guessing, mostly wrongly — reliably beat learners who spend the same total time studying. The failed attempt is not a cost of assessment. It is a teaching event in its own right. The literature now calls its benefit the pretesting effect.

The errorless inheritance

To see how heretical this finding once was, recall where the fear of errors came from. Mid-century behaviorism treated learning as the strengthening of stimulus–response bonds. That made every error a rehearsal of the wrong bond — something to engineer out of instruction entirely. Teaching machines were built to make mistakes nearly impossible. They stepped learners through material in steps so small that the correct response was almost guaranteed.

The doctrine outlived the theory that produced it. Long after cognitive psychology had replaced behaviorism, instructional design kept its reflexes. Explain thoroughly. Practice gently. Treat a learner’s error as a failure of the explanation.

The cognitive account of memory tells a different story. Suppose learning is the building and strengthening of retrieval routes, not the stamping-in of responses. Then an error is not a rehearsal of the wrong answer. It is evidence that a search happened — and the search itself changes the searcher. The pretesting literature supplied the demonstration. This is not just possible in theory; it is reliably true, under conditions ordinary courses can reproduce.

Guessing wrong on purpose

The cleanest lab version used weakly linked word pairs, such as whale–mammal. One group studied each pair for the full trial. The other saw only whale–? first, tried to produce the partner word, almost always failed, and then saw the answer for the remaining seconds. Despite spending less time with the correct answer visible, the guess-first group recalled substantially more on the final test (Kornell, Hays & Bjork, 2009). The paper’s title states the finding without hedging: unsuccessful retrieval attempts enhance subsequent learning.

The design detail that makes these experiments persuasive is the control condition. The guess-first group is not compared against people who did nothing. It is compared against people who spent the same total time — every second the pretest group used on failing, the study group used on studying. If pretesting were merely a way of spending time near the material, the conditions would tie. They do not tie, repeatedly. That forces the conclusion: the attempt itself has teaching value that studying the same content for the same seconds does not (Kornell, Hays & Bjork, 2009).

The effect is not a quirk of word lists. Some students took a pretest on questions covered by an essay they had not yet read — and failed, of course. After reading, they remembered the relevant material better than students given the same extra time to study the essay itself (Richland, Kornell & Kao, 2009). With video lectures, brief prequestions improved memory for the prequestioned content in later viewing (Carpenter & Toftness, 2017). Reviews pooling dozens of experiments find the effect robust across materials and test formats, with the largest gains on exactly the content the pretest targeted (Pan & Carpenter, 2023).

The number

Dozens of experiments, aggregated in review, find the guess-first advantage robust across materials and test formats — with the largest gains on exactly the content the pretest targeted (Pan & Carpenter, 2023).

CONDITION A · STUDY ONLY CONDITION B · PRETEST, THEN STUDY study the material attempt ✗ study the material final test final test ▲ Same total time. The failed attempt wins the final test — a reliable advantage across word pairs, prose, and video. © 2026 FUTURE PROOF™
Figure 1. The standard pretesting design. Time is equated; only the sequence differs. Attempting first — and failing — beats pure study on the delayed test. Design schematic after Kornell, Hays & Bjork (2009) and Richland et al. (2009). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Why failure primes learning

A motivational channel runs alongside the cognitive mechanisms taken up below, and it is not decorative. An attempted-and-missed question turns the upcoming answer from information into resolution. The learner now has a specific open loop the material will close. And self-reports in errorful-learning studies consistently describe heightened curiosity about the correct answer after a committed guess (Metcalfe, 2017). Course designers spend real money trying to manufacture “engagement” with narrative and gamification. A question the learner just failed generates the genuine article for free — aimed at exactly the content the module needs them to want.

Three mechanisms carry most of the explanatory weight. First, activation: searching memory for an answer you don’t have wakes the neighborhood of related knowledge, giving the incoming answer a richer structure to attach to (Grimaldi & Karpicke, 2012). Second, gap awareness: the failed attempt turns “content to be covered” into “a specific question I couldn’t answer,” and that changes how the later material is read. Third, error correction itself: far from stamping in mistakes, errors followed by feedback are potent learning events (Metcalfe, 2017).

The counterintuitive capstone is the hypercorrection effect. Errors made with high confidence are the ones most likely to be corrected and retained (Butterfield & Metcalfe, 2001). Why? Apparently because the surprise of being wrong commands attention.

Careful comparisons show the benefit is not mere exposure to the question format. Generating an error beats being shown the same question with its answer, and beats reading an elaborated version of the material (Potts & Shanks, 2014). Something about the attempt itself does the work.

What counts as an attempt

Because the attempt is the active ingredient, its quality matters. The literature is fairly specific about what a real attempt involves. It must be committed: the learner produces or selects an answer before anything is revealed, rather than glancing at a question and mentally shrugging toward the answer key. It should involve search. The benefit is strongest when the learner has related knowledge to draw on. That is why guessing at word pairs related in meaning helps, while guessing at arbitrary pairings helps far less (Grimaldi & Karpicke, 2012).

And it works best when the learner briefly believes the question is answerable. A prompt that reads as an absurd demand invites disengagement. A prompt that reads as a fair challenge — one a competent person might meet — recruits genuine retrieval effort. With it come the activation and curiosity that make the next answer stick.

Format matters less than commitment, but it is not nothing. Open-ended attempts force fuller search than multiple-choice recognition. And multiple-choice pretests carry a specific hazard the open format avoids: plausible wrong options can themselves gain familiarity. The practical rule the evidence supports has three parts. Prefer generation where the learner has anything to generate from. Use recognition formats when the material is wholly new — and in either case make the correction unmissable, which is where the design question hands off to the feedback literature.

Chance the error is corrected after feedback hypercorrection what the studies find what people assume more often less often most correctedlow confidence high confidence confidence in the wrong answer © 2026 FUTURE PROOF™
Figure 2. The hypercorrection effect, plotted against the intuition it overturns. The dashed grey line is what almost everyone assumes: that the more firmly a learner holds a wrong answer, the harder it is to dislodge. The teal line is what the studies find — the chance a wrong answer is corrected and retained after feedback rises with how sure the learner was of it, so the coral point at the far right, confidently wrong, is the one most likely to be fixed (Butterfield & Metcalfe, 2001). The coral bracket is the distance between the two accounts at maximum confidence. The surprise of being wrong recruits the attention the correction rides on, which is why the worst moment in a learner’s session is the best teaching opportunity it will produce. Both lines are ordinal illustrations, not measured values — read the direction of each slope and the gap between them, not the plotted heights. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Unsuccessful retrieval attempts enhance subsequent learning. The title — and the finding — of Kornell, Hays & Bjork (2009).

The pretest as an instrument, not just an intervention

Everything above treats the pretest as a treatment applied to the learner. Its second life — arguably its more valuable one at scale — is as a measurement applied to the curriculum. A pretest given before a module is a map of prior knowledge, drawn at exactly the moment the map can still change what happens next.

Read the items. Items most of a cohort already answers correctly flag content the module can compress or skip. Items nobody gets flag where the module’s real work lies. Items that split the cohort flag where one fixed lesson will bore half the room while losing the other half. Instructors have always known this in principle. What changed is that a system delivering the pretest can act on the map at once, per learner, rather than averaging it into next year’s course revision.

The error patterns are even richer than the scores. Say two learners both miss a question about statistical power. They are not in the same state if one confused power with significance while the other confused it with sample size. The wrong answers name the specific misconception; a simple miss only names a gap.

Analysing which wrong option each learner chose is, in effect, a misconception census taken before teaching begins. It changes what “personalized” can mean: not just faster or slower through the same slides, but different explanations aimed at different documented confusions. None of this telemetry exists in a course that opens with content. That is among the quieter arguments for asking first. The questions are free to move, and in their pre-instruction position they earn twice.

Pretesting vs. posttesting

Anyone who has read our testing-effect review will ask the natural question. If quizzing after study works, and quizzing before study works, which is better? Direct comparisons suggest the two effects are similar in size and partly complementary (Pan & Sana, 2021). Pretesting shapes how new material is taken in and stored; posttesting strengthens retrieval of material already stored. A curriculum has no reason to choose. The evidence-based sequence is: attempt, study, retrieve — questions on both sides of the content.

Composed this way, the two effects turn a module into a loop rather than a lecture. The opening attempt primes the encoding and hands the system a map of what this learner is missing. The study phase arrives as the answer to live questions instead of an unprovoked broadcast. The closing retrieval cements what was just built and generates the next round of telemetry. Each question does double duty — teaching the learner and informing the system — and that is precisely what makes question-first design economical rather than merely virtuous. The same item bank powers diagnosis, instruction, and reinforcement; only its position in the sequence changes what it does.

The feedback condition

One boundary matters more than all the others: the answer must arrive. The pretesting literature is unanimous that the benefit depends on corrective feedback following the attempt. Errors left uncorrected can persist. And timing studies show the correction does its work when the learner still cares about the question (Hays, Kornell & Bjork, 2013). Pretesting without prompt feedback is not a desirable difficulty — it is just difficulty.

The catch

The whole effect is conditional on the answer arriving. Deliver the correction in the same screen flow, while the question is still warm — not behind an optional review link at the module’s end — or the error you deliberately invited is free to persist.

The hypercorrection finding sharpens the point into something almost paradoxical. The worst moment in a learner’s session — confidently wrong — is the best teaching opportunity the session will produce (Butterfield & Metcalfe, 2001). A confident error means the learner has a structured, committed belief for the correction to collide with. The surprise recruits attention no headline or animation can buy. Systems that hide wrong answers to spare feelings, or bury the correction behind an optional “review” link, are discarding the most valuable events they generate. The design rule is the opposite of gentleness-by-omission: surface the miss plainly, explain it at once, and treat the learner’s surprise as the mechanism, not a malfunction.

What the evidence doesn’t show

  • It is not a license to make assessments punishing. The experiments use low-stakes attempts where failure carries no cost. Graded pretests would import test anxiety into exactly the moment the technique needs playful engagement.
  • The benefit concentrates on pretested content. Prequestions most reliably help memory for the material they target; effects on untested neighboring content are smaller and less consistent (Carpenter & Toftness, 2017), (Pan & Carpenter, 2023). Pretests are aiming devices, not general accelerants — which is an argument for generating them broadly across the syllabus.
  • Most delays are short. The bulk of the laboratory evidence tests retention within a session to a few days. The long-horizon durability of pretesting gains, unlike spacing’s, is still being mapped.
  • Learners will not thank you at first. As with every desirable difficulty, the experience of failing feels inefficient, and learners rate errorless study as more effective even when their own scores disagree. The interface has to carry the explanation.

Where the evidence stops

  1. 1It is not a license to make assessments punishing
  2. 2The benefit concentrates on pretested content
  3. 3Most delays are short
  4. 4Learners will not thank you at first
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

Put questions before content, everywhere the content is new. The change costs almost nothing — the questions already exist in the post-quiz. Moving even a handful of them to the front of a module converts passive opening minutes into the highest-leverage part of the session. Keep the attempts short, visibly low-stakes, and framed honestly. Learners should be told, in as many words, that they are expected to miss most of these — and that missing is the point. The framing matters because the technique’s only real enemy is the learner’s inference that failure means the course is broken, or they are.

Deliver the answer, with an explanation, while the question is still warm — in the same screen flow, not in a summary at the module’s end. Target pretests at the ideas the module most needs to land, because the benefit concentrates on what was asked. Prefer generation over recognition wherever the learner has anything to generate from.

And treat wrong answers as the most valuable telemetry a lesson can produce. Each one is a documented gap, timestamped at the moment before instruction arrived to fill it. Each confident wrong answer is a flagged priority. A course that collects and acts on its learners’ errors is running the pretesting literature as infrastructure. A course that merely tolerates them leaves the effect’s second half unclaimed.

Applied research

How Future Proof™ applies this: diagnostic-first modules.

Every Future Proof module opens with adaptive probe questions, not a video. Learners attempt, mostly miss, and get the explanation while the question is still live — the pretesting sequence, wired into the flow. The engine then does what a laboratory pretest cannot: it keeps the errors. Each wrong answer updates the learner’s knowledge map, steers which content is served next, and seeds the spaced-review schedule — so the failed attempt teaches twice.

See the adaptive diagnostic
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 2001Butterfield
  • 2009Kornell
  • 2009Richland
  • 2012Grimaldi
  • 2013Hays
  • 2014Potts
  • 2017Carpenter
  • 2017Metcalfe
  • 2021Pan
  • 2023Pan
© 2026 FUTURE PROOF™
The evidence base. The 10 sources cited here span 2001–2023, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Kornell, N., Hays, M.J., & Bjork, R.A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition 35(4): 989–998. PDF
  2. Richland, L.E., Kornell, N., & Kao, L.S. (2009). The pretesting effect: Do unsuccessful retrieval attempts enhance learning? Journal of Experimental Psychology: Applied 15(3): 243–257. PDF
  3. Carpenter, S.K., & Toftness, A.R. (2017). The effect of prequestions on learning from video presentations. Journal of Applied Research in Memory and Cognition 6(1): 104–109. PDF
  4. Pan, S.C., & Carpenter, S.K. (2023). Prequestioning and pretesting effects: A review of empirical research, theoretical perspectives, and implications for practice. Educational Psychology Review 35: 97. DOI
  5. Grimaldi, P.J., & Karpicke, J.D. (2012). When and why do retrieval attempts enhance subsequent encoding? Memory & Cognition 40(4): 505–513. PDF
  6. Metcalfe, J. (2017). Learning from errors. Annual Review of Psychology 68: 465–489. DOI
  7. Butterfield, B., & Metcalfe, J. (2001). Errors committed with high confidence are hypercorrected. Journal of Experimental Psychology: Learning, Memory, and Cognition 27(6): 1491–1494. PDF
  8. Potts, R., & Shanks, D.R. (2014). The benefit of generating errors during learning. Journal of Experimental Psychology: General 143(2): 644–667. PDF
  9. Pan, S.C., & Sana, F. (2021). Pretesting versus posttesting: Comparing the pedagogical benefits of errorful generation and retrieval practice. Journal of Experimental Psychology: Applied 27(2): 237–257. PDF
  10. Hays, M.J., Kornell, N., & Bjork, R.A. (2013). When and why a failed test potentiates the effectiveness of subsequent study. Journal of Experimental Psychology: Learning, Memory, and Cognition 39(1): 290–296. PDF
Try the AI engine

See what your learners can’t answer yet.

Book a 20-minute demo with your team’s actual content. We’ll show you the diagnostic-first flow — and how the errors it collects become each learner’s personal curriculum.

10 citations Reviewed August 2026 Open peer review welcomed