© 2026 FUTURE PROOF™
Memory & Practice · Desirable Difficulties

Desirable Difficulties: Why Easy Training Fails

Conditions that slow visible learning — spacing, interleaving, testing — reliably deepen it. A tour of Robert Bjork’s research programme, the fluency illusion it exposed, and how Future Proof™ turns productive struggle into a design principle.

TL;DR

The finding: Conditions of practice that slow or complicate acquisition — spacing sessions apart, interleaving topics, testing instead of re-presenting — depress performance during training and improve retention and transfer after it. Conditions that make training feel smooth do the reverse. Performance while learning is a systematically misleading index of what was learned.

The mechanism: What you observe during training reflects momentary retrieval strength — how accessible the material is right now. Learning is growth in storage strength, and that growth is largest precisely when retrieval is effortful. Fluent practice inflates the first without moving the second, so learners, instructors, and dashboards all overestimate what stuck.

The product: Future Proof’s adaptive difficulty engine keeps each learner in productive struggle instead of fluent cruising, and retake sets are reshuffled and re-sampled so a second attempt measures the skill — not a memorized answer key.

In this article

  1. 01Learning is not performance
  2. 02The difficulties that earn the adjective
  3. 03Fluency is a misleading signal
  4. 04Storage strength, retrieval strength
  5. 05Designing with two strengths in mind
  6. 06What the evidence doesn’t show
  7. 07What this means for training design
© 2026 FUTURE PROOF™
The route. 7 sections, from “Learning is not performance” to “What this means for training design”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Most of this library’s memory cluster — spacing, testing, interleaving, pretesting — consists of specific techniques with specific evidence. This article is about the framework that unifies them. It covers the single uncomfortable principle from which each technique falls out as a special case, and the illusion that explains why every one of them is chronically underused.

Every training programme generates two kinds of evidence about itself. The first is visible right away: quiz scores climbing across a module, exercises finished faster, error rates thinning by the session. The second arrives months later, when somebody has to use the skill under real conditions. The uncomfortable finding at the centre of Robert Bjork’s research programme is that these two signals routinely disagree. Worse, the conditions of practice that optimise the first often degrade the second (Soderstrom & Bjork, 2015).

Bjork named the phenomenon in a chapter on the training of human beings: desirable difficulties (Bjork, 1994). These are changes to practice that slow the visible rate of learning and add errors and effort in the moment. Yet they reliably improve retention and transfer measured weeks or months later. Both words are load-bearing. The difficulties are genuine: learners perform worse during training, and they tend to resent them. And they are desirable only under conditions the same literature is careful to spell out.

Learning is not performance

The framework rests on a distinction that sounds pedantic and turns out to be the whole game. Performance is what can be observed and scored during practice — accuracy on the practice set, speed on the current block. Learning is the relatively permanent change in knowledge or skill that shows up on delayed tests of retention and transfer. Soderstrom and Bjork’s integrative review assembles decades of evidence that the two come apart, in both directions. Learning can accumulate while measured performance barely moves. And performance can soar during training while leaving little durable change behind (Soderstrom & Bjork, 2015).

The cleanest demonstrations came out of motor-skills research before spreading to verbal learning. Schmidt and Bjork reviewed experiments across both domains and found a recurring crossover. Variable practice, random task ordering, and less frequent feedback all depressed the learning curve — and improved performance on tests given after a delay (Schmidt & Bjork, 1992). An instructor choosing between schedules on the strength of the practice curve alone would have picked, in case after case, the one that produced less learning.

This is the trap Bjork described for training organisations three decades ago (Bjork, 1994). The measures available while training is under way — momentary accuracy, speed, learner confidence — are not neutral gauges. They lean toward whatever schedule makes retrieval easy right now. And that is precisely the schedule that leaves the shallowest trace.

during practice delayed test blocked · massed · restudied spaced · interleaved · tested retention gap Performance © 2026 FUTURE PROOF™
Figure 1. The acquisition–retention crossover. The condition that looks better during practice (red, dashed) loses to the more difficult condition (purple, solid) once the test is delayed. Schematic illustration of the pattern reviewed in Soderstrom & Bjork (2015); axes are illustrative, not measured data. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The difficulties that earn the adjective

“Desirable difficulties” is a family label, and each member has its own evidence base.

Spacing. Spreading practice across time, rather than massing it into one block, is among the oldest and best-replicated moves in experimental psychology. Cepeda and colleagues pooled several hundred experimental comparisons. Spaced practice beat massed practice on delayed retention tests across ages, materials, and intervals — even though massing often looks as good or better when the test comes right away (Cepeda et al., 2006).

Retrieval practice. Testing yourself is a difficulty in exactly Bjork’s sense. Roediger and Karpicke had students either restudy a prose passage or take a recall test on it. On a test given minutes later, the restudy group did better. But when the final test came days later, the ordering reversed, and the tested group retained substantially more (Roediger & Karpicke, 2006). The crossover is the signature of the whole literature — the condition that felt worse and scored worse early produced the stronger memory.

Interleaving. Mixing problem types during practice, instead of blocking them by category, is the difficulty learners resent most. Rohrer and Taylor gave students mathematics practice that was either blocked or shuffled. The blocked group did far better during practice, and the shuffled group did far better on a test a week later (Rohrer & Taylor, 2007). Kornell and Bjork found the same pattern for learning categories from examples. People learned to identify painters’ styles better when examples of different artists were interleaved — even though most believed the massed presentation had worked better for them (Kornell & Bjork, 2008).

Variation and leaner feedback. Varying the conditions of practice, and giving feedback less often or only in summary form, both slow early progress. Both also improve delayed retention and transfer in the motor-learning literature (Schmidt & Bjork, 1992).

Bjork and Bjork’s later chapter draws these threads into a single account written for non-specialists. It remains the best short statement of the programme (Bjork & Bjork, 2011).

The primary goal of instruction should be to facilitate long-term learning — that is, to create relatively permanent changes in comprehension, understanding, and skills of the types that will support long-term retention and transfer. Soderstrom & Bjork 2015, Learning Versus Performance

Fluency is a misleading signal

A framework that indicts nearly all teaching intuition owes an account of how the intuition went wrong. The answer turns out to be a measurable illusion with its own experimental literature, and it operates on learners and institutions alike. Why do learners, instructors, and institutions keep choosing the schedules that lose? Because the felt experience of learning tracks ease of processing, not durable storage. When material comes to mind fluently — because it was just presented, because the answer is in view, because the practice block repeats a single skill — that fluency is read as knowledge.

Koriat and Bjork called the result an illusion of competence. Learners judged how well they would later remember material while the answer sat in front of them — and were markedly overconfident relative to their actual recall (Koriat & Bjork, 2005). The illusion survives contact with the evidence. In Kornell and Bjork’s interleaving study, most people still rated the massed presentation as more effective after the interleaved condition had just outperformed it for them personally (Kornell & Bjork, 2008). Simon and Bjork found the motor-skill version. Learners practising under a blocked schedule predicted better future performance than those under a random schedule — and the prediction was backwards (Simon & Bjork, 2001).

The catch

The illusion survives contact with its own refutation: participants kept rating the massed schedule as more effective immediately after the interleaved one had outperformed it for them personally. You cannot ask learners which schedule worked — the feeling of fluency answers first, and it answers wrong.

The consequences show up in what learners choose to do with their time. Dunlosky and colleagues rated ten common study techniques in a landmark monograph. The strategies learners lean on most — re-reading and highlighting — ranked among the least effective, while practice testing and spaced practice earned the highest utility ratings (Dunlosky et al., 2013). Fluent techniques feel productive; effective ones feel like struggle. And any dashboard that reports in-training performance inherits the same bias, at institutional scale.

The number

10 techniques Rated head-to-head in the Dunlosky monograph: the two learners use most — re-reading and highlighting — landed among the least effective, while practice testing and distributed practice took the top utility ratings (Dunlosky et al., 2013).

Storage strength, retrieval strength

A collection of crossover effects invites a unifying account — some model of memory under which spacing, testing, interleaving, and lean feedback all fall out of one principle rather than four coincidences. The field has one, and its two variables repay learning by name. The standard account is Bjork and Bjork’s “new theory of disuse” (Bjork & Bjork, 1992). On this model every memory carries two strengths. Retrieval strength is how accessible the item is right now; storage strength is how deeply it is entrenched. Performance during training reflects retrieval strength; learning, in the sense that matters, is growth in storage strength.

The model’s central asymmetry: the gain in storage strength from a successful retrieval is larger when retrieval strength is low — that is, when access is effortful. Practising something while it is still easy to reach buys little. Letting it become difficult before retrieving it buys much more. Easy training, on this account, is not a gentler version of the real thing. It is practice spent where it buys least — comfort purchased at the price of the durability the training existed to create.

Designing with two strengths in mind

The two-strength model is unusually useful as a design instrument, because it converts vague intentions into checkable questions. For any practice activity, ask: what is the learner’s retrieval strength for this material right now, and is this activity scheduled where the storage-strength purchase is large? A review served minutes after first exposure fails the check — retrieval strength is still high, so the session buys little. The same review after days of forgetting passes it. A quiz whose answers remain visible fails it outright, since nothing is retrieved at all (Koriat & Bjork, 2005). Run across a curriculum, the check moves minutes out of the fluent activities that dominate course design, and into the effortful ones where the model says learning is actually transacted (Bjork & Bjork, 1992).

The model also dictates the measurement reform this article keeps circling. If in-training performance reflects retrieval strength, then every dashboard built on immediate scores is reporting the wrong strength. Optimizing against it — as renewal decisions and A/B tests do — actively selects for shallow designs. The fix is structural, not a matter of attitude. Make the delayed assessment the assessment of record, and let session metrics inform pacing while being barred from verdicts (Soderstrom & Bjork, 2015). Organizations that make this one change discover that their course rankings reshuffle — and that the new ranking finally agrees with what their best instructors always suspected.

Design rule

Make the delayed assessment the assessment of record. Session metrics may inform pacing, but any score collected at the moment of peak accessibility is barred from verdicts — because it is measuring the wrong strength, and optimizing against it selects for shallow designs.

Finally, the model prescribes the learner-communication work that most deployments skip. Difficulties feel like malfunction to the people experiencing them. The metacognition studies show the feeling does not correct itself with experience (Kornell & Bjork, 2008), (Simon & Bjork, 2001). Telling learners the design is deliberate — naming the crossover, showing them their own delayed results — is not a courtesy but a retention feature. The struggle that learners understand as mechanism is struggle they will tolerate long enough for it to pay.

What the evidence doesn’t show

The literature is often compressed to “harder is better,” a slogan its own authors explicitly reject. Several limits deserve equal billing.

Difficulty is desirable only when the learner can overcome it. Bjork and Bjork are explicit that the same manipulation becomes an undesirable difficulty for a learner who lacks the background knowledge or skills to respond to it successfully (Bjork & Bjork, 2011). Struggle that ends in failure with no corrective feedback is just failure.

There is no dosing formula. The studies identify which side of the easy–hard trade to err on. They do not supply a validated universal setpoint for how much difficulty is best, and the right level plainly varies with the learner and the material. Claims of a single “optimal difficulty number” outrun the evidence.

Support varies by technique and setting. Dunlosky and colleagues rated practice testing and spacing highly on classroom-relevant evidence. They rated interleaving lower precisely because much of its support at the time came from lab studies of narrow materials (Dunlosky et al., 2013). And much of the founding work uses retention gaps of days to weeks, not the months that workplace training cares about. That is one reason Soderstrom and Bjork argue that evaluations of training should be based on delayed, not immediate, assessments (Soderstrom & Bjork, 2015).

Not every difficulty qualifies. The label covers a specific set of manipulations whose benefits run through retrieval effort and encoding variability. It is not a licence to make training arbitrarily frustrating and call the frustration pedagogy.

Practice testing Distributed practice Re-reading Highlighting highest utility highest utility low utility — what learners lean on most low utilityUtility ratings of common study techniques (Dunlosky et al., 2013) © 2026 FUTURE PROOF™
Figure 2. The ranking that indicts the default habits: practice testing and distributed practice earned the monograph’s highest utility ratings, while re-reading and highlighting — the techniques learners lean on most — landed among the least effective. Bar lengths schematic after Dunlosky et al. (2013); read the ordering, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Applied in practice

How Future Proof™ applies this.

Two mechanisms in the platform are direct implementations of this literature. The adaptive difficulty engine selects each next item near the edge of the learner’s current ability, so practice stays in productive struggle instead of fluent cruising — when accuracy gets comfortable, the questions get harder, and when struggle stops being productive, they ease off. And retakes are reshuffled: a second attempt draws a re-ordered, re-sampled question set rather than repeating the original items, so improvement reflects the skill — not a memorized answer key.

See the adaptive engine

What this means for training design

For anyone building or buying training, the implications are blunt:

  • Evaluate on a delay. End-of-course scores measure retrieval strength at its peak. If the outcome you care about is performance months later, the assessment has to happen later — or on transfer tasks the training never showed.
  • Treat smoothness as ambiguous. A cohort cruising through a module at high accuracy is compatible with excellent teaching and with practice pitched too easy to leave a trace. In-training comfort is not evidence of learning.
  • Engineer the difficulties in. Space sessions apart, interleave related topics, and put retrieval — not re-presentation — at the centre of practice. These are scheduling decisions, which means they can be made by design rather than left to learner preference, which reliably runs the other way.
  • Protect assessments from memorisation. A retake that repeats the same items in the same order measures memory for the answer key, not the skill. Parallel or reshuffled forms keep retrieval, rather than recognition, doing the work.

Where the evidence stops

  1. 1Evaluate on a delay
  2. 2Treat smoothness as ambiguous
  3. 3Engineer the difficulties in
  4. 4Protect assessments from memorisation
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
References

Selected papers.

These are the studies cited above — not an exhaustive bibliography of the desirable-difficulties literature. DOI links go to the publisher’s version; where no stable DOI exists, the link searches Google Scholar.

The evidence, by year

  • 1992Schmidt
  • 1992Bjork
  • 1994Bjork
  • 2001Simon
  • 2005Koriat
  • 2006Cepeda
  • 2006Roediger
  • 2007Rohrer
  • 2008Kornell
  • 2011Bjork
  • 2013Dunlosky
  • 2015Soderstrom
© 2026 FUTURE PROOF™
The evidence base. The 12 sources cited here span 1992–2015, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Soderstrom, N.C., & Bjork, R.A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science 10(2): 176–199. DOI
  2. Schmidt, R.A., & Bjork, R.A. (1992). New conceptualizations of practice: Common principles in three paradigms suggest new concepts for training. Psychological Science 3(4): 207–217. PDF
  3. Bjork, R.A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing About Knowing, pp. 185–205. MIT Press. PDF
  4. Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132(3): 354–380. DOI
  5. Roediger, H.L., & Karpicke, J.D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science 17(3): 249–255. DOI
  6. Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science 35(6): 481–498. DOI
  7. Kornell, N., & Bjork, R.A. (2008). Learning concepts and categories: Is spacing the “enemy of induction”? Psychological Science 19(6): 585–592. DOI
  8. Bjork, E.L., & Bjork, R.A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M.A. Gernsbacher et al. (Eds.), Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society, pp. 56–64. Worth Publishers. PDF
  9. Koriat, A., & Bjork, R.A. (2005). Illusions of competence in monitoring one’s knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition 31(2): 187–194. DOI
  10. Simon, D.A., & Bjork, R.A. (2001). Metacognition in motor learning. Journal of Experimental Psychology: Learning, Memory, and Cognition 27(4): 907–912. DOI
  11. Dunlosky, J., Rawson, K.A., Marsh, E.J., Nathan, M.J., & Willingham, D.T. (2013). Improving students’ learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest 14(1): 4–58. DOI
  12. Bjork, R.A., & Bjork, E.L. (1992). A new theory of disuse and an old theory of stimulus fluctuation. In A. Healy, S. Kosslyn, & R. Shiffrin (Eds.), From Learning Processes to Cognitive Processes: Essays in Honor of William K. Estes, Vol. 2, pp. 35–67. Erlbaum. PDF
See it in practice

Productive struggle, by design.

Book a 20-minute demo with your team’s actual content. We’ll show you how the difficulty curve adapts for one real learner — and put a reshuffled retake next to the original so you can see what it protects against.

12 citations Reviewed August 2026 Open peer review welcomed