Research · Memory & Practice
Memory & Practice · Interleaving

Interleaving: the counterintuitive ordering effect.

Blocked practice feels productive and tests poorly; mixed practice feels clumsy and wins the delayed test. What the discrimination-learning literature says about the order of practice — and how Future Proof™’s confusion-pair coaching puts it to work.

TL;DR

The finding: When practice on related-but-distinct topics is mixed together (interleaved) rather than grouped by topic (blocked), performance during practice gets worse — and performance on a delayed test gets better, often dramatically. The result holds for mathematics procedures and for learning categories, and it survives meta-analysis.

The mechanism: Blocked practice lets learners execute a solution without ever choosing it, and it hides the differences between confusable ideas. Interleaving forces discrimination: every item first asks which kind of problem is this? before what do I do about it?

The product: Future Proof’s confusion-pair coaching deliberately interleaves look-alike concepts a learner has previously mixed up — turning the learner’s own error history into a discrimination-training schedule.

Open almost any textbook, course, or training module and you will find the same architecture: a lesson on one technique, followed by a set of exercises that all use that technique. Practice one thing until it feels smooth, then move to the next. This is blocked practice, and it is how nearly all instruction is organized — because it feels efficient to learners and looks efficient to instructors.

The alternative sounds almost careless. Shuffle problems from several related topics together, so that consecutive items rarely call for the same approach. Learners under this interleaved schedule make more errors during practice, report the sessions as harder, and rate the method as less effective. Then they sit a delayed test — and they win, often by a wide margin. That reversal between what practice feels like and what it produces is one of the most instructive findings in learning science, because it explains why so much well-intentioned curriculum design is organized backwards.

A shuffled problem set, a reversed result

The cleanest demonstration comes from a study in which college students learned to compute the volumes of four obscure geometric solids (Rohrer & Taylor, 2007). Every student solved the same practice problems; the only manipulation was order. One group worked in blocks — all problems of one type, then all of the next — while the other worked through the same problems shuffled together.

During practice, the blocked group was far more accurate. That is unsurprising: each problem repeated the procedure just used, so there was nothing to figure out. One week later, on a test of novel problems, the ordering reversed sharply — the interleaved group answered 63% correctly against 20% for the blocked group (Rohrer & Taylor, 2007).

The interpretation the authors offered has held up well since. Under blocking, a problem effectively announces its own solution: if the last ten problems were wedge problems, the eleventh is too. Students practice executing a procedure but never practice selecting one. A test — like every real situation in which knowledge is used — does not announce which technique applies. Interleaved practice is practice at choosing; blocked practice is practice at executing a choice someone else already made.

100% 75% 50% 25% 89% 60% During practice 20% 63% Test, one week later blocked interleaved Accuracy
Figure 1. Accuracy during practice vs. one week later under blocked (red) and interleaved (purple) schedules. The schedule that looks better while it happens loses on the delayed test. Approximate values redrawn from Rohrer & Taylor (2007).

Twelve painters and a stubborn illusion

Interleaving is not just for procedures. In a now-classic study of inductive learning, participants studied landscape paintings by twelve artists, either blocked — several works by one artist in a row — or interleaved, with the artists shuffled together (Kornell & Bjork, 2008). The test was to identify the artists of paintings never seen before.

The expectation, shared by the researchers and by nearly everyone shown the design, was that blocking should help: seeing an artist’s works side by side ought to let their common style emerge. The data said otherwise. Interleaved study produced reliably better classification of new paintings. The paper’s title asked whether spacing is the “enemy of induction”; the data answered that it is closer to a friend.

The most consequential part of the study, though, was metacognitive. Asked afterwards which schedule had helped them more, a large majority of participants said blocking — even after their own test scores had demonstrated the opposite (Kornell & Bjork, 2008). The schedule that works loses the exit survey.

Is spacing the “enemy of induction”? The title question of Kornell & Bjork (2008). Their participants said yes. Their data said no.

The mechanism: learning to tell things apart

Why would a worse-feeling schedule produce better learning? The dominant account is discriminative contrast. Much of what we call “understanding a concept” is really the ability to tell it apart from its confusable neighbors — this solid from that one, this artist from a similar one, this clause type from the one it resembles. Blocked practice emphasizes what members of a category share; interleaving juxtaposes categories, so their differences become salient exactly when they can be compared.

Several lines of evidence support this reading. The interleaving advantage in style-learning appears specifically when the temporal arrangement promotes contrast between categories, rather than spacing per se (Kang & Pashler, 2012). Work isolating the components suggests both between-category discrimination and retrieval from long-term memory contribute (Birnbaum, Kornell, Bjork & Bjork, 2013). And the benefit is not unconditional: when categories are easy to tell apart, blocking can be competitive, with the advantage of interleaving concentrated on categories that are genuinely confusable (Carvalho & Goldstone, 2014).

Interleaving also inevitably brings spacing along with it — mixing topics spreads each topic’s repetitions out in time. Studies designed to separate the two contributions in mathematics learning find that discriminative contrast and distributed practice both do real work (Foster, Mueller, Was, Rawson & Dunlosky, 2019). In practice this is a feature, not a confound: a well-built interleaved schedule delivers two evidence-backed effects at once.

The meta-analytic picture

The literature has been aggregated. A meta-analysis of the interleaving experiments found an overall benefit of interleaved over blocked practice that is small-to-medium on average (Hedges’ g ≈ 0.42) — but heavily moderated by material (Brunmair & Richter, 2019). The benefit is large for confusable visual categories such as paintings and for mathematics problem types, and it shrinks toward zero — in some materials reversing — for content like foreign-language vocabulary, where the items studied do not compete with one another (Brunmair & Richter, 2019).

The meta-analysis is titled “Similarity matters,” and that is the correct summary. Interleaving is not a universal accelerant; it is a discrimination-training tool. It pays when the learner must tell similar things apart, and it pays most when those things are most easily confused.

From laboratory to classroom

The effect scales beyond the lab. Fourth graders who practiced mixed problem types scored roughly double their blocked-practice peers on a next-day test (Taylor & Rohrer, 2010). With middle-school mathematics students, the interleaving advantage was larger on a test given thirty days after practice than one day after — the gap grows with delay (Rohrer, Dedrick & Stershic, 2015). And a large preregistered randomized controlled trial across dozens of seventh-grade classrooms found a large advantage for interleaved worksheets on a delayed test (d ≈ 0.83) (Rohrer, Dedrick, Hartwig & Cheung, 2020).

Set that against how instructional materials are actually built. Textbooks block by chapter; courses block by module; drill software blocks by skill. The ordering that the classroom evidence favors is the one almost no published curriculum uses.

Why blocked practice fools learners and instructors

The deeper lesson is about measurement. Performance during learning is a poor proxy for learning — and in this literature it is an inverted one. Under blocking, practice is fast, fluent, and light on errors, and learners read that fluency as mastery. Instructors watching the same session see smooth progress and satisfied students, and draw the same wrong conclusion.

The illusion is remarkably resistant to correction. Even when participants experience both schedules and see their own interleaved-condition scores come out ahead, most continue to believe blocking worked better, misattributing the fluency of blocked study to superior learning (Yan, Bjork & Bjork, 2016). Both sides of the classroom, in other words, prefer the schedule that produces less durable learning — which is why the correction has to be built into the system that sequences practice, not left to judgment in the moment.

What the evidence doesn’t show

Interleaving has earned its place in the evidence-based toolkit, but the literature is specific about its boundaries, and honest applications should be too.

  • It is not “mix anything with anything.” The benefit depends on the interleaved items being confusable. Alternating algebra with history trains no discrimination, and meta-analytic estimates for low-similarity materials such as foreign vocabulary are near zero or negative (Brunmair & Richter, 2019).
  • Blocking is not always wrong. When categories are easy to distinguish, blocked study can match or beat interleaving (Carvalho & Goldstone, 2014), and hybrid schedules — a short blocked introduction before mixing — remain under-studied relative to how often they are recommended.
  • The long-horizon evidence is concentrated in mathematics. The strongest classroom trials involve school math over delays of days to a month (Rohrer, Dedrick, Hartwig & Cheung, 2020). Extension to workplace skills and multi-month retention is plausible but rests on a thinner base.
  • Interleaving and spacing travel together. Many designs cannot fully separate the two, and where they are separated, both mechanisms contribute (Foster et al., 2019) — so “interleaving” gains reported in applied settings are usually a package, not a single ingredient.

What this means for practice

Treat sequence as an instructional variable with the same status as content. Shuffle practice within families of confusable problem types, not across unrelated subjects. Expect — and tell learners to expect — lower accuracy during practice, because the dip is the mechanism working, not a defect. Evaluate on delayed tests rather than end-of-session scores, since session performance rewards exactly the wrong schedule. And treat a learner’s specific confusions as a map: the pairs of ideas an individual actually mixes up are the highest-value candidates for interleaved contrast.

Applied research

How Future Proof™ applies this: confusion-pair coaching.

When a learner keeps mixing up two look-alike concepts — precision and recall, accrual and deferral, two clauses of the same regulation — the platform logs the pair from their answer history. Confusion-pair coaching then does what the discrimination literature prescribes: it deliberately interleaves those exact concepts in later review sessions, posing questions that force the learner to tell them apart rather than recognize either in isolation. The mix is rebuilt from each learner’s own errors, so the interleaving is aimed at precisely the discriminations that learner still needs.

See confusion-pair coaching
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

  1. Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science 35(6): 481–498. DOI
  2. Kornell, N., & Bjork, R.A. (2008). Learning concepts and categories: Is spacing the “enemy of induction”? Psychological Science 19(6): 585–592. DOI
  3. Kang, S.H.K., & Pashler, H. (2012). Learning painting styles: Spacing is advantageous when it promotes discriminative contrast. Applied Cognitive Psychology 26(1): 97–103. PDF
  4. Birnbaum, M.S., Kornell, N., Bjork, E.L., & Bjork, R.A. (2013). Why interleaving enhances inductive learning: The roles of discrimination and retrieval. Memory & Cognition 41(3): 392–402. PDF
  5. Carvalho, P.F., & Goldstone, R.L. (2014). Putting category learning in order: Category structure and temporal arrangement affect the benefit of interleaved over blocked study. Memory & Cognition 42(3): 481–495. PDF
  6. Foster, N.L., Mueller, M.L., Was, C., Rawson, K.A., & Dunlosky, J. (2019). Why does interleaving improve math learning? The contributions of discriminative contrast and distributed practice. Memory & Cognition 47(6): 1088–1101. PDF
  7. Brunmair, M., & Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin 145(11): 1029–1052. DOI
  8. Taylor, K., & Rohrer, D. (2010). The effects of interleaved practice. Applied Cognitive Psychology 24(6): 837–848. DOI
  9. Rohrer, D., Dedrick, R.F., & Stershic, S. (2015). Interleaved practice improves mathematics learning. Journal of Educational Psychology 107(3): 900–908. DOI
  10. Rohrer, D., Dedrick, R.F., Hartwig, M.K., & Cheung, C.-N. (2020). A randomized controlled trial of interleaved mathematics practice. Journal of Educational Psychology 112(1): 40–52. PDF
  11. Yan, V.X., Bjork, E.L., & Bjork, R.A. (2016). On the difficulty of mending metacognitive illusions: A priori theories, fluency effects, and misattributions of the interleaving benefit. Journal of Experimental Psychology: General 145(7): 918–933. PDF
Try the AI engine

See what your learners keep mixing up.

Book a 20-minute demo with your team’s actual content. We’ll show you the confusion pairs Future Proof detects in real answer data — and the interleaved review schedule it builds from them.

11 citations Reviewed July 2026 Open peer review welcomed