Research · The Uncomfortable Evidence
The Uncomfortable Evidence · Gamification

Gamification: what the meta-analyses say.

Streaks, points and leaderboards are not magic and not snake oil. They lift some outcomes reliably, leave others untouched, and quietly corrode a few — and which happens depends on the design. Here is the honest meta-analytic read, and where Future Proof™ chooses to spend the mechanic.

TL;DR

The finding: Across meta-analyses, gamification produces a small-to-moderate positive effect on learning — largest for cognitive outcomes, smaller and less stable for motivation and behaviour (Sailer & Homner, 2020). But the average hides enormous variance: the same points-and-leaderboards kit that raises effort in one study leaves intrinsic motivation flat in another (Mekler et al., 2017). The honest headline is not “it works” — it’s “it depends, and we can say fairly precisely on what.”

The mechanism: Game elements do not carry learning by themselves. They change how much and how a person engages, and that engagement only pays off when it is pointed at an action that produces learning. Aimed at retrieval, feedback and clear goals, the mechanic amplifies. Aimed at superficial busywork — or bolted onto an already-loved task as a controlling reward — it can flatten or even undermine (Deci, Koestner & Ryan, 1999).

The product: On Future Proof, streaks, XP and team leaderboards are tied to retrieval practice — the gamified action is the evidence-backed action, so the mechanic amplifies the thing that already builds durable memory rather than a proxy for it.

Ask a product team whether adding a streak counter, a points balance and a leaderboard will make people learn more, and you will get one of two confident answers. The optimist points to Duolingo. The cynic points to every abandoned corporate LMS with a dusty “badges” tab. The literature agrees with neither, because the honest answer is conditional — and the conditions are, by now, reasonably well mapped.

The first serious synthesis set the tone. Hamari, Koivisto and Sarsa reviewed the empirical gamification studies available by 2014 and reached a verdict that has aged well: gamification generally produces positive effects, but those effects depend heavily on the context in which it is deployed and on the people using it (Hamari, Koivisto & Sarsa, 2014). They also flagged a problem that still haunts the field — many early studies measured self-reported enjoyment rather than learning, and confounded the game elements with the novelty of a new system.

A decade later, we have proper meta-analyses instead of narrative reviews. They sharpen the picture without overturning it.

What the meta-analyses actually found

The most cited quantitative synthesis in learning contexts is Sailer and Homner’s 2020 meta-analysis in Educational Psychology Review. Pooling controlled comparisons, they report significant but modest effects that differ by outcome type: a small-to-moderate effect on cognitive learning outcomes (g ≈ 0.49), a smaller effect on motivational outcomes (g ≈ 0.36), and a smaller one again on behavioural outcomes (g ≈ 0.25) (Sailer & Homner, 2020). The pattern matters as much as the numbers. Gamification’s clearest win is on learning, not — as the folk theory would have it — on motivation.

Two details keep that finding honest. First, when the authors restricted the analysis to studies with high methodological rigour, the cognitive effect held up while the motivational and behavioural effects became less stable (Sailer & Homner, 2020). The result you would most expect gamification to produce — a motivation bump — is the one that wobbles most under scrutiny. Second, a separate meta-analysis focused on education reached a compatible conclusion, an overall medium-sized advantage for gamified over non-gamified instruction (g ≈ 0.50), while its authors documented, from the qualitative record, both why learners enjoy these systems and why some come to resent them (Bai, Hew & Huang, 2020).

So the field’s average is real and positive. The trouble — and the interesting part — is the spread around it.

0 0.25 0.50 0.75 Effect size (g) vs non-gamified instruction Cognitive outcomes (Sailer & Homner 2020) 0.49 Education overall (Bai et al. 2020) 0.50 Motivational outcomes (Sailer & Homner 2020) 0.36 Behavioural outcomes (Sailer & Homner 2020) 0.25 Points/levels/leaderboards on intrinsic motivation ≈ null (Mekler et al. 2017)
Figure 1. Gamification’s effect is largest on cognitive learning and smallest — and least stable — on motivation and behaviour. Points, levels and leaderboards, isolated, moved performance but not intrinsic motivation. Effect sizes are Hedges’ g; outcome definitions and comparison conditions differ across studies.

The element matters more than the label

“Gamification” is not one thing, and treating it as one is the single biggest source of confused findings. The most informative studies stop asking whether gamification works and start asking which element does what.

Mekler, Brühlmann, Tuch and Opwis ran the cleanest test of the canonical trio. In a controlled experiment on an image-annotation task, they compared points, levels and leaderboards against a plain control. The result is quietly important: the game elements significantly increased performance — people did more — but did not significantly affect intrinsic motivation or perceived competence (Mekler et al., 2017). In other words, these mechanics behave like extrinsic incentives that direct effort, not like magic that makes a boring task feel meaningful. As Mekler and colleagues put it, points, levels and leaderboards by themselves neither make nor break intrinsic motivation in a non-game context.

Sailer and colleagues came at it from the other side, asking which needs each element feeds. Using a self-determination-theory lens, they found that badges, leaderboards and performance graphs bolstered the experience of competence and of meaningful goals, while socially oriented elements — avatars, a narrative, teammates — did more for the sense of relatedness (Sailer et al., 2017). The lesson is not that one set is better; it is that different mechanics pull different psychological levers, so the right element depends on which lever the task needs.

Leaderboards are the clearest case of a mechanic whose value is conditional on why it works. Landers, Bauer and Callan showed that a leaderboard raised task performance about as much as assigning people a specific, difficult goal did — and that the effect ran through goal-setting, moderated by how committed each person was to the goal (Landers, Bauer & Callan, 2017). Read carefully, that is a warning as much as an endorsement: a leaderboard helps to the extent it functions as a clear, committed-to goal. Strip out the goal and leave only the social comparison, and there is no reason to expect the same lift.

The effects [of gamification] are greatly dependent on the context in which the gamification is being implemented, as well as on the users using it. Hamari, Koivisto & Sarsa, HICSS 2014

When it backfires: the overjustification risk

The most important cautionary evidence predates gamification entirely. In a meta-analysis of 128 experiments, Deci, Koestner and Ryan found that tangible, expected rewards contingent on doing a task reliably undermined people’s later free-choice engagement with that task — a negative effect on intrinsic motivation, strongest precisely for the reward structures gamification loves to imitate (Deci, Koestner & Ryan, 1999). The mechanism, in self-determination terms, is that a salient controlling reward shifts a person’s felt reason for acting from “I want to” to “I’m being paid to,” and the intrinsic reason erodes.

This is the overjustification trap, and gamification walks straight toward it whenever it staples points onto something people already found worthwhile. The practical corollary is uncomfortable for anyone selling “engagement”: the more genuinely interesting a task already is, the more a crude points overlay risks cheapening it. Where gamification is safest is exactly where intrinsic interest is lowest — the repetitive, effortful, easy-to-skip work that learning actually requires. Which is a useful clue about where to aim it.

What the evidence doesn’t show

It is easy to over-read a positive meta-analytic average. Four things this literature does not establish:

  • It does not show a large, dependable effect. The pooled effects are small-to-moderate, and the motivational and behavioural ones weaken under high-rigour analysis (Sailer & Homner, 2020). Anyone quoting gamification as a reliable multiplier is quoting the top of a wide, skewed distribution.
  • It does not show that the game elements are doing the work. Many primary studies confound the mechanic with novelty, with more time-on-task, or with better feedback introduced alongside it. Reviews have repeatedly flagged short durations and the risk that measured gains are a novelty effect that fades (Hamari, Koivisto & Sarsa, 2014).
  • It does not show a motivation boost you can bank on. The canonical points/levels/leaderboards trio moved performance but not intrinsic motivation in a clean experiment (Mekler et al., 2017), and under the wrong framing, contingent rewards can push intrinsic motivation the other way (Deci, Koestner & Ryan, 1999).
  • It does not show the effect is uniform across people or contexts. Both the framing reviews and the moderator analyses find that outcomes vary with the user, the task and the competitive structure, and the volume of genuinely mixed results is itself a headline finding (Hamari, Koivisto & Sarsa, 2014). Averages here conceal more than they reveal.

None of this makes gamification worthless. It relocates the question. The right question is never “does gamification work?” — it is “which element, aimed at which action, for which learner, measured against which outcome?”

The design rule that follows

Put the strands together and a single design principle falls out. Game mechanics reliably do one thing: they increase and direct effort. They do not, on their own, make that effort productive, and clumsily applied they can taint tasks that were fine without them. So the leverage is entirely in what you attach the mechanic to. Attach a streak, a point or a leaderboard to a low-value action — logging in, watching a video to completion — and you get more of a low-value action. Attach it to an action the learning science already endorses, and the mechanic borrows that action’s validity.

The evidence-backed action, across almost every review of durable learning, is retrieval practice: being made to recall, not merely re-encounter. That is the action worth gamifying, because there the extra effort the mechanic buys is spent on the one behaviour most tightly linked to remembering.

Applied at Future Proof

How Future Proof™ applies this — gamify the evidence-backed action.

We take the meta-analytic caveats literally, so every mechanic is bolted to retrieval, never to a proxy for it. Streaks count consecutive days of actually recalling material — not logins, not video-completions — so the habit the streak protects is the habit that builds memory. XP is earned by answering retrieval questions and clearing mastery checks, which keeps the points pointed at learning outcomes rather than time-on-task. And team leaderboards are framed as committed, specific goals in the goal-setting sense the evidence supports, and scoped to teams to blunt the intrinsic-motivation risk that raw individual ranking carries. The mechanic amplifies; the retrieval does the learning.

See how the mechanics work
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Research Library PDF.

  1. Hamari, J., Koivisto, J., & Sarsa, H. (2014). Does gamification work? — A literature review of empirical studies on gamification. Proceedings of the 47th Hawaii International Conference on System Sciences (HICSS): 3025–3034. DOIPDF
  2. Sailer, M., & Homner, L. (2020). The gamification of learning: a meta-analysis. Educational Psychology Review 32(1): 77–112. DOI
  3. Mekler, E.D., Brühlmann, F., Tuch, A.N., & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior 71: 525–534. DOIPDF
  4. Deci, E.L., Koestner, R., & Ryan, R.M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin 125(6): 627–668. DOI
  5. Bai, S., Hew, K.F., & Huang, B. (2020). Does gamification improve student learning outcome? Evidence from a meta-analysis and synthesis of qualitative data in educational contexts. Educational Research Review 30: 100322. DOI
  6. Sailer, M., Hense, J.U., Mayr, S.K., & Mandl, H. (2017). How gamification motivates: An experimental study of the effects of specific game design elements on psychological need satisfaction. Computers in Human Behavior 69: 371–380. DOI
  7. Landers, R.N., Bauer, K.N., & Callan, R.C. (2017). Gamification of task performance with leaderboards: A goal setting experiment. Computers in Human Behavior 71: 508–515. DOI
  8. Koivisto, J., & Hamari, J. (2019). The rise of motivational information systems: A review of gamification research. International Journal of Information Management 45: 191–210. DOI
See the mechanics

Gamify the action the evidence endorses — not a proxy for it.

Book a 20-minute demo using your team’s actual content. We’ll show you where the streaks, XP and leaderboards attach — always to retrieval practice — and why that keeps the game amplifying learning instead of gaming a metric.

8 citations Reviewed July 2026 Open peer review welcomed