Research · Memory & Practice
Memory & Practice · Desirable Difficulties

Desirable Difficulties: Why Easy Training Fails

Conditions that slow visible learning — spacing, interleaving, testing — reliably deepen it. A tour of Robert Bjork’s research programme, the fluency illusion it exposed, and how Future Proof™ turns productive struggle into a design principle.

TL;DR

The finding: Conditions of practice that slow or complicate acquisition — spacing sessions apart, interleaving topics, testing instead of re-presenting — depress performance during training and improve retention and transfer after it. Conditions that make training feel smooth do the reverse. Performance while learning is a systematically misleading index of what was learned.

The mechanism: What you observe during training reflects momentary retrieval strength — how accessible the material is right now. Learning is growth in storage strength, and that growth is largest precisely when retrieval is effortful. Fluent practice inflates the first without moving the second, so learners, instructors, and dashboards all overestimate what stuck.

The product: Future Proof’s adaptive difficulty engine keeps each learner in productive struggle instead of fluent cruising, and retake sets are reshuffled and re-sampled so a second attempt measures the skill — not a memorized answer key.

Every training programme generates two kinds of evidence about itself. The first is visible immediately: quiz scores climbing across a module, exercises finished faster, error rates thinning by the session. The second arrives months later, when somebody has to use the skill under real conditions. The uncomfortable finding at the centre of Robert Bjork’s research programme is that these two signals routinely disagree — and that the conditions of practice that optimise the first often degrade the second (Soderstrom & Bjork, 2015).

Bjork named the phenomenon in a chapter on the training of human beings: desirable difficulties (Bjork, 1994) — manipulations of practice that slow the visible rate of acquisition, add errors and effort in the moment, and yet reliably improve retention and transfer measured weeks or months later. Both words are load-bearing. The difficulties are genuine: learners perform worse during training, and they tend to resent them. And they are desirable only under conditions the same literature is careful to specify.

Learning is not performance

The framework rests on a distinction that sounds pedantic and turns out to be the whole game. Performance is what can be observed and scored during acquisition — accuracy on the practice set, speed on the current block. Learning is the relatively permanent change in knowledge or skill that shows up on delayed tests of retention and transfer. Soderstrom and Bjork’s integrative review assembles decades of evidence that the two dissociate in both directions: learning can accumulate while measured performance barely moves, and performance can soar during training while leaving little durable change behind (Soderstrom & Bjork, 2015).

The cleanest demonstrations came out of motor-skills research before spreading to verbal learning. Schmidt and Bjork reviewed experiments across both domains and found a recurring crossover: manipulations such as variable practice, random task ordering, and less frequent feedback depressed the acquisition curve — and improved performance on tests given after a delay (Schmidt & Bjork, 1992). An instructor choosing between schedules on the basis of the acquisition curve alone would have picked, in case after case, the one that produced less learning.

This is the trap Bjork described for training organisations three decades ago (Bjork, 1994). The measures available while training is under way — momentary accuracy, speed, learner confidence — are not neutral indicators. They are biased toward whatever schedule makes retrieval easy right now, which is precisely the schedule that leaves the shallowest trace.

during practice delayed test blocked · massed · restudied spaced · interleaved · tested retention gap Performance
Figure 1. The acquisition–retention crossover. The condition that looks better during practice (red, dashed) loses to the more difficult condition (purple, solid) once the test is delayed. Schematic illustration of the pattern reviewed in Soderstrom & Bjork (2015); axes are illustrative, not measured data.

The difficulties that earn the adjective

“Desirable difficulties” is a family label, and each member has its own evidence base.

Spacing. Distributing practice across time, rather than massing it into one block, is among the oldest and best-replicated manipulations in experimental psychology. Cepeda and colleagues’ quantitative synthesis of several hundred experimental comparisons found that spaced practice beat massed practice on delayed retention tests across ages, materials, and intervals — even though massing frequently looks as good or better when the test comes immediately (Cepeda et al., 2006).

Retrieval practice. Testing yourself is a difficulty in exactly Bjork’s sense. Roediger and Karpicke had students either restudy a prose passage or take a recall test on it: on a test given minutes later, the restudy group did better, but when the final test came days later the ordering reversed and the tested group retained substantially more (Roediger & Karpicke, 2006). The crossover is the signature of the whole literature — the condition that felt worse and scored worse early produced the stronger memory.

Interleaving. Mixing problem types during practice, instead of blocking them by category, is the difficulty learners find most aversive. Rohrer and Taylor gave students mathematics practice that was either blocked or shuffled: the blocked group performed far better during practice, and the shuffled group performed far better on a test a week later (Rohrer & Taylor, 2007). Kornell and Bjork found the same pattern for inductive learning — participants learned to identify painters’ styles better when examples of different artists were interleaved, even though most participants believed the massed presentation had worked better for them (Kornell & Bjork, 2008).

Variation and leaner feedback. Varying the conditions of practice, and providing feedback less often or only in summary form, both slow acquisition and improve delayed retention and transfer in the motor-learning literature (Schmidt & Bjork, 1992).

Bjork and Bjork’s later chapter draws these threads into a single account written for non-specialists, and it remains the best short statement of the programme (Bjork & Bjork, 2011).

The primary goal of instruction should be to facilitate long-term learning — that is, to create relatively permanent changes in comprehension, understanding, and skills of the types that will support long-term retention and transfer. Soderstrom & Bjork 2015, Learning Versus Performance

Fluency is a misleading signal

Why do learners, instructors, and institutions keep choosing the schedules that lose? Because the subjective experience of learning tracks ease of processing, not durable storage. When material comes to mind fluently — because it was just presented, because the answer is in view, because the practice block repeats a single skill — that fluency is read as knowledge.

Koriat and Bjork called the result an illusion of competence: learners who judged how well they would later remember material while the answer was in front of them were markedly overconfident relative to their actual recall (Koriat & Bjork, 2005). The illusion survives contact with the evidence. In Kornell and Bjork’s interleaving study, most participants still rated the massed presentation as more effective after the interleaved condition had just outperformed it for them personally (Kornell & Bjork, 2008). Simon and Bjork found the motor-skill version: learners practising under a blocked schedule predicted better future performance than those practising under a random schedule, and the prediction was backwards (Simon & Bjork, 2001).

The consequences show up in what learners choose to do with their time. Dunlosky and colleagues’ monograph rating ten common study techniques found the strategies learners lean on most — re-reading and highlighting — among the least effective, while practice testing and distributed practice earned the highest utility ratings (Dunlosky et al., 2013). Fluent techniques feel productive; effective ones feel like struggle. And any dashboard that reports in-training performance inherits the same bias, at institutional scale.

Storage strength, retrieval strength

The standard theoretical account is Bjork and Bjork’s “new theory of disuse” (Bjork & Bjork, 1992). On this model every memory carries two strengths: retrieval strength, how accessible the item is right now, and storage strength, how deeply it is entrenched. Performance during training reflects retrieval strength; learning, in the sense that matters, is growth in storage strength. The model’s central asymmetry is that the gain in storage strength from a successful retrieval is larger when retrieval strength is low — that is, when access is effortful. Practising something while it is still easy to reach purchases little; letting it become difficult before retrieving it purchases much more. Easy training, on this account, is not a gentler version of the real thing. It is practice spent where it buys least.

What the evidence doesn’t show

The literature is often compressed to “harder is better,” a slogan its own authors explicitly reject. Several limits deserve equal billing.

Difficulty is desirable only when the learner can overcome it. Bjork and Bjork are explicit that the same manipulation becomes an undesirable difficulty for a learner who lacks the background knowledge or skills to respond to it successfully (Bjork & Bjork, 2011). Struggle that ends in failure with no corrective feedback is just failure.

There is no dosing formula. The studies identify which side of the easy–hard trade to err on; they do not supply a validated universal setpoint for how much difficulty is optimal, and the right level plainly varies with the learner and the material. Claims of a single “optimal difficulty number” outrun the evidence.

Support varies by technique and setting. Dunlosky and colleagues rated practice testing and spacing highly on classroom-relevant evidence, but rated interleaving lower precisely because much of its support at the time came from laboratory studies of narrow materials (Dunlosky et al., 2013). And much of the foundational work uses retention intervals of days to weeks rather than the months that workplace training cares about — one reason Soderstrom and Bjork argue that evaluations of training should be based on delayed, not immediate, assessments (Soderstrom & Bjork, 2015).

Not every difficulty qualifies. The label covers a specific set of manipulations whose benefits run through retrieval effort and encoding variability. It is not a licence to make training arbitrarily frustrating and call the frustration pedagogy.

Applied in practice

How Future Proof™ applies this.

Two mechanisms in the platform are direct implementations of this literature. The adaptive difficulty engine selects each next item near the edge of the learner’s current ability, so practice stays in productive struggle instead of fluent cruising — when accuracy gets comfortable, the questions get harder, and when struggle stops being productive, they ease off. And retakes are reshuffled: a second attempt draws a re-ordered, re-sampled question set rather than repeating the original items, so improvement reflects the skill — not a memorized answer key.

See the adaptive engine

What this means for training design

For anyone building or buying training, the implications are blunt:

  • Evaluate on a delay. End-of-course scores measure retrieval strength at its peak. If the outcome you care about is performance months later, the assessment has to happen later — or on transfer tasks the training never showed.
  • Treat smoothness as ambiguous. A cohort cruising through a module at high accuracy is compatible with excellent teaching and with practice pitched too easy to leave a trace. In-training comfort is not evidence of learning.
  • Engineer the difficulties in. Space sessions apart, interleave related topics, and put retrieval — not re-presentation — at the centre of practice. These are scheduling decisions, which means they can be made by design rather than left to learner preference, which reliably runs the other way.
  • Protect assessments from memorisation. A retake that repeats the same items in the same order measures memory for the answer key, not the skill. Parallel or reshuffled forms keep retrieval, rather than recognition, doing the work.
References

Selected papers.

These are the studies cited above — not an exhaustive bibliography of the desirable-difficulties literature. DOI links go to the publisher’s version; where no stable DOI exists, the link searches Google Scholar.

  1. Soderstrom, N.C., & Bjork, R.A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science 10(2): 176–199. DOI
  2. Schmidt, R.A., & Bjork, R.A. (1992). New conceptualizations of practice: Common principles in three paradigms suggest new concepts for training. Psychological Science 3(4): 207–217. PDF
  3. Bjork, R.A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing About Knowing, pp. 185–205. MIT Press. PDF
  4. Cepeda, N.J., Pashler, H., Vul, E., Wixted, J.T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132(3): 354–380. DOI
  5. Roediger, H.L., & Karpicke, J.D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science 17(3): 249–255. DOI
  6. Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science 35(6): 481–498. DOI
  7. Kornell, N., & Bjork, R.A. (2008). Learning concepts and categories: Is spacing the “enemy of induction”? Psychological Science 19(6): 585–592. DOI
  8. Bjork, E.L., & Bjork, R.A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M.A. Gernsbacher et al. (Eds.), Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society, pp. 56–64. Worth Publishers. PDF
  9. Koriat, A., & Bjork, R.A. (2005). Illusions of competence in monitoring one’s knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition 31(2): 187–194. DOI
  10. Simon, D.A., & Bjork, R.A. (2001). Metacognition in motor learning. Journal of Experimental Psychology: Learning, Memory, and Cognition 27(4): 907–912. DOI
  11. Dunlosky, J., Rawson, K.A., Marsh, E.J., Nathan, M.J., & Willingham, D.T. (2013). Improving students’ learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest 14(1): 4–58. DOI
  12. Bjork, R.A., & Bjork, E.L. (1992). A new theory of disuse and an old theory of stimulus fluctuation. In A. Healy, S. Kosslyn, & R. Shiffrin (Eds.), From Learning Processes to Cognitive Processes: Essays in Honor of William K. Estes, Vol. 2, pp. 35–67. Erlbaum. PDF
See it in practice

Productive struggle, by design.

Book a 20-minute demo with your team’s actual content. We’ll show you how the difficulty curve adapts for one real learner — and put a reshuffled retake next to the original so you can see what it protects against.

12 citations Reviewed July 2026 Open peer review welcomed