© 2026 FUTURE PROOF™
The Uncomfortable Evidence · Gamification

Gamification: what the meta-analyses say.

Streaks, points and leaderboards are not magic and not snake oil. They lift some outcomes reliably, leave others untouched, and quietly corrode a few — and which happens depends on the design. Here is the honest meta-analytic read, and where Future Proof™ chooses to spend the mechanic.

TL;DR

The finding: Across meta-analyses, gamification produces a small-to-moderate positive effect on learning — largest for cognitive outcomes, smaller and less stable for motivation and behaviour (Sailer & Homner, 2020). But the average hides enormous variance: the same points-and-leaderboards kit that raises effort in one study leaves intrinsic motivation flat in another (Mekler et al., 2017). The honest headline is not “it works” — it’s “it depends, and we can say fairly precisely on what.”

The mechanism: Game elements do not carry learning by themselves. They change how much and how a person engages, and that engagement only pays off when it is pointed at an action that produces learning. Aimed at retrieval, feedback and clear goals, the mechanic amplifies. Aimed at superficial busywork — or bolted onto an already-loved task as a controlling reward — it can flatten or even undermine (Deci, Koestner & Ryan, 1999).

The product: On Future Proof, streaks, XP and team leaderboards are tied to retrieval practice — the gamified action is the evidence-backed action, so the mechanic amplifies the thing that already builds durable memory rather than a proxy for it.

In this article

  1. 01What the meta-analyses actually found
  2. 02How the field graded its own homework
  3. 03The element matters more than the label
  4. 04When it backfires: the overjustification risk
  5. 05What the evidence doesn’t show
  6. 06The design rule that follows
© 2026 FUTURE PROOF™
The route. 6 sections, from “What the meta-analyses actually found” to “The design rule that follows”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Few design fashions have travelled as far on as thin a diet of evidence as gamification did in its first decade. The word was coined around 2010, and the conference keynotes followed within a year. By the time the first systematic review appeared, points and badges had already been bolted onto banking apps, fitness trackers, and most of the world’s learning platforms. The research has spent the years since catching up with the rollout. The catch-up produced something more useful than a verdict: a fairly precise map of which mechanics move which outcomes, under which conditions, for which people. This article is that map.

Ask a product team whether adding a streak counter, a points balance and a leaderboard will make people learn more, and you will get one of two confident answers. The optimist points to Duolingo. The cynic points to every abandoned corporate LMS with a dusty “badges” tab. The literature agrees with neither. The honest answer is conditional — and the conditions are, by now, fairly well mapped.

The first serious synthesis set the tone. Hamari, Koivisto and Sarsa reviewed the empirical studies available by 2014 and reached a verdict that has aged well. Gamification generally produces positive effects. But those effects depend heavily on the context and on the people using it (Hamari, Koivisto & Sarsa, 2014). They also flagged a problem that still haunts the field. Many early studies measured self-reported enjoyment rather than learning, and could not separate the game elements from the novelty of a new system.

A decade later, we have proper meta-analyses — studies that pool the results of many studies — instead of narrative reviews. They sharpen the picture without overturning it.

What the meta-analyses actually found

Meta-analysis matters more in this literature than in most, because the single studies disagree so much. Individual results range from strongly positive to null to negative. Only pooling shows whether the centre of that spread sits above zero. And only the moderator analyses — the breakdowns by condition — show what drives the spread.

The most cited quantitative synthesis in learning contexts is Sailer and Homner’s 2020 meta-analysis in Educational Psychology Review. Pooling controlled comparisons, they report significant but modest effects that differ by outcome type. Cognitive learning outcomes: a small-to-moderate effect (g ≈ 0.49). Motivation outcomes: smaller (g ≈ 0.36). Behaviour outcomes: smaller again (g ≈ 0.25) (Sailer & Homner, 2020). The pattern matters as much as the numbers: gamification’s clearest win is on learning, not — as the folk theory would have it — on motivation.

Two details keep that finding honest. First, the authors re-ran the analysis on only the most rigorous studies. The learning effect held up; the motivation and behaviour effects became less stable (Sailer & Homner, 2020). The result you would most expect gamification to produce — a motivation bump — is the one that wobbles most under scrutiny. Second, a separate meta-analysis focused on education reached a matching conclusion: an overall medium-sized advantage for gamified over non-gamified instruction (g ≈ 0.50). Its authors also documented, from learners’ own accounts, both why people enjoy these systems and why some come to resent them (Bai, Hew & Huang, 2020).

So the field’s average is real and positive. The trouble — and the interesting part — is the spread around it. A wide spread around a modest mean guarantees that many real deployments sit at zero or below.

0 0.25 0.50 0.75 Effect size (g) vs non-gamified instruction Cognitive outcomes (Sailer & Homner 2020) 0.49 Education overall (Bai et al. 2020) 0.50 Motivational outcomes (Sailer & Homner 2020) 0.36 Behavioural outcomes (Sailer & Homner 2020) 0.25 Points/levels/leaderboards on intrinsic motivation ≈ null (Mekler et al. 2017) © 2026 FUTURE PROOF™
Figure 1. Gamification’s effect is largest on cognitive learning and smallest — and least stable — on motivation and behaviour. Points, levels and leaderboards, isolated, moved performance but not intrinsic motivation. Effect sizes are Hedges’ g; outcome definitions and comparison conditions differ across studies. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

How the field graded its own homework

Between the early narrative review and the mature meta-analyses sits a document worth knowing about. Koivisto and Hamari’s 2019 review read several hundred empirical studies of the by-then-sprawling field, and it graded the field’s methods as much as its findings (Koivisto & Hamari, 2019). Their census confirmed education and learning as gamification’s main home.

Their sharper contribution was a catalogue of the field’s recurring weaknesses. Studies too short to separate real effects from novelty. Missing control conditions. Self-report used where behaviour could have been measured. And a persistent habit of studying “gamification” as a bundle rather than isolating which element did what.

That catalogue is the right lens for every effect size in this article. It explains why the rigorous-subset analyses matter more than the headline averages. It explains why single-element experiments like Mekler’s carry weight beyond their size. And it explains why anyone reading a glowing case study should ask two questions first: how long did it run, and what was the comparison? A field that knows its own failure modes this precisely is a field worth trusting — about exactly as far as its better studies go.

The catch

Before believing any glowing gamification case study, ask the field’s own two questions: how long did it run, and what was the comparison? Short studies cannot separate the mechanic from novelty, and a missing control cannot separate it from anything at all.

The element matters more than the label

“Gamification” is not one thing. Treating it as one is the single biggest source of confused findings. The most useful studies stop asking whether gamification works and start asking which element does what.

Mekler, Brühlmann, Tuch and Opwis ran the cleanest test of the classic trio. In a controlled experiment on an image-annotation task, they compared points, levels and leaderboards against a plain control. The result is quietly important. The game elements significantly increased performance — people did more — but did not significantly affect intrinsic motivation or perceived competence (Mekler et al., 2017). In other words, these mechanics act like external incentives that direct effort — not like magic that makes a boring task feel meaningful. As Mekler and colleagues put it, points, levels and leaderboards by themselves neither make nor break intrinsic motivation in a non-game context.

Sailer and colleagues came at it from the other side. They asked which needs each element feeds, using the lens of self-determination theory — psychology’s standard account of what feeds motivation. Badges, leaderboards and performance graphs strengthened the sense of competence and of meaningful goals. Social elements — avatars, a narrative, teammates — did more for the sense of relatedness (Sailer et al., 2017). The lesson is not that one set is better. Different mechanics pull different psychological levers, so the right element depends on which lever the task needs.

Leaderboards are the clearest case of a mechanic whose value depends on why it works. Landers, Bauer and Callan showed that a leaderboard raised task performance about as much as assigning people a specific, difficult goal did — and that the effect ran through goal-setting, hinging on how committed each person was to the goal (Landers, Bauer & Callan, 2017). Read carefully, that is a warning as much as an endorsement. A leaderboard helps to the extent it works as a clear goal people commit to. Strip out the goal and leave only the social comparison, and there is no reason to expect the same lift.

The effects [of gamification] are greatly dependent on the context in which the gamification is being implemented, as well as on the users using it. Hamari, Koivisto & Sarsa, HICSS 2014

When it backfires: the overjustification risk

Everything so far has treated the worst case as zero — a mechanic that does nothing. The motivation literature contains a worse worst case. It is old enough and large enough that no gamification design should be signed off without checking against it.

The most important warning evidence predates gamification entirely. In a meta-analysis of 128 experiments, Deci, Koestner and Ryan studied tangible, expected rewards tied to doing a task. Such rewards reliably undermined people’s later free-choice engagement with that task (Deci, Koestner & Ryan, 1999). It is a negative effect on intrinsic motivation, strongest precisely for the reward structures gamification loves to copy. The mechanism, in self-determination terms: a salient controlling reward shifts a person’s felt reason for acting from “I want to” to “I’m being paid to.” The intrinsic reason erodes.

The number

128 experiments The meta-analytic base of the overjustification caution: tangible, expected rewards contingent on doing a task reliably undermined later free-choice engagement with it (Deci, Koestner & Ryan, 1999) — strongest for exactly the structures gamification imitates.

This is the overjustification trap. Gamification walks straight toward it whenever it staples points onto something people already found worthwhile. The practical upshot is uncomfortable for anyone selling “engagement”: the more genuinely interesting a task already is, the more a crude points overlay risks cheapening it. Where gamification is safest is exactly where intrinsic interest is lowest — the repetitive, effortful, easy-to-skip work that learning actually requires. Which is a useful clue about where to aim it.

Two boundary conditions keep the caution in proportion. Deci and colleagues found the undermining effect concentrated in tangible, expected, contingent rewards — rewards tied to doing the task. Verbal feedback and informational signals of competence did not carry the same risk, and could even boost intrinsic motivation (Deci, Koestner & Ryan, 1999). That distinction is the designer’s escape route. A mechanic that reads as information about progress — a mastery bar, a performance graph — sits on the safe side of the line; one that reads as payment for compliance sits on the dangerous side. The same badge can be either, depending on framing. The trap is real, but it is a trap with a posted map.

What the evidence doesn’t show

It is easy to over-read a positive meta-analytic average. Four things this literature does not establish:

  • It does not show a large, dependable effect. The pooled effects are small-to-moderate, and the motivational and behavioural ones weaken under high-rigour analysis (Sailer & Homner, 2020). Anyone quoting gamification as a reliable multiplier is quoting the top of a wide, skewed distribution.
  • It does not show that the game elements are doing the work. Many primary studies confound the mechanic with novelty, with more time-on-task, or with better feedback introduced alongside it. Reviews have repeatedly flagged short durations and the risk that measured gains are a novelty effect that fades (Hamari, Koivisto & Sarsa, 2014).
  • It does not show a motivation boost you can bank on. The canonical points/levels/leaderboards trio moved performance but not intrinsic motivation in a clean experiment (Mekler et al., 2017), and under the wrong framing, contingent rewards can push intrinsic motivation the other way (Deci, Koestner & Ryan, 1999).
  • It does not show the effect is uniform across people or contexts. Both the framing reviews and the moderator analyses find that outcomes vary with the user, the task and the competitive structure, and the volume of genuinely mixed results is itself a headline finding (Hamari, Koivisto & Sarsa, 2014). Averages here conceal more than they reveal.

Where the evidence stops

  1. 1It does not show a large, dependable effect
  2. 2It does not show that the game elements are doing the work
  3. 3It does not show a motivation boost you can bank on
  4. 4It does not show the effect is uniform across people or contexts
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

None of this makes gamification worthless. It moves the question. The right question is never “does gamification work?” — it is “which element, aimed at which action, for which learner, measured against which outcome?” Asked that way, the literature stops being a referendum and becomes a design manual — one whose chapters happen to be effect sizes.

The design rule that follows

Put the strands together — the outcome-level averages, the element-level experiments, the needs mapping, the overjustification boundary — and a single design principle falls out. Game mechanics reliably do one thing: they increase and direct effort. They do not, on their own, make that effort productive — and clumsily applied, they can taint tasks that were fine without them. So the leverage sits entirely in what you attach the mechanic to. Attach a streak, a point or a leaderboard to a low-value action — logging in, watching a video to the end — and you get more of a low-value action. Attach it to an action the learning science already endorses, and the mechanic borrows that action’s validity.

Design rule

The leverage is entirely in the attachment point. Before shipping any mechanic, name the exact action it rewards — if that action is not one the learning evidence endorses on its own, the mechanic is buying more of the wrong thing, efficiently.

Across almost every review of durable learning, the evidence-backed action is retrieval practice: being made to recall, not merely see the material again. That is the action worth gamifying. There, the extra effort the mechanic buys is spent on the one behaviour most tightly linked to remembering.

The same principle sorts the common mechanics into an honest ranking. A streak that counts days of retrieval protects a habit worth having; a streak that counts logins protects a metric. XP earned per correct recall keeps the incentive pointed at the outcome; XP earned per minute of video watched pays people to leave the tab open. A leaderboard framed as a specific, committed team goal borrows the goal-setting evidence (Landers, Bauer & Callan, 2017); a raw individual ranking borrows only the social-comparison risks. In each pair the mechanic is identical. The difference is entirely in the action it is soldered to — which is why two products with the same feature list can sit at opposite ends of the effect spread.

And the overjustification evidence supplies the final filter: aim mechanics at the tasks people avoid, not the ones they already enjoy (Deci, Koestner & Ryan, 1999). Daily retrieval of half-forgotten material is effortful, unglamorous, and chronically skipped. That makes it at once the highest-value target in learning and the safest one in motivation terms. Gamification’s best use, it turns out, is not making learning fun. It is making the unfun part of learning get done.

no change Tangible, expected, contingent Badge read as payment Points / levels / leaderboards Badge read as information Verbal / informational feedback≈ no shift ← undermines enhances → Later intrinsic motivation (ordinal) © 2026 FUTURE PROOF™
Figure 2. The overjustification boundary as a diverging chart: bars run left where a mechanic tends to erode later intrinsic motivation, right where it can strengthen it. Across 128 experiments, tangible rewards that were expected and contingent on doing the task reliably undermined later free-choice engagement, while verbal and informational signals of competence carried no such risk and could enhance it (Deci, Koestner & Ryan, 1999). The isolated points, levels and leaderboards trio moved performance but left intrinsic motivation unmoved (Mekler et al., 2017). The same badge lands on either side of the line depending on whether it reads as payment or as information about progress. Bar lengths are ordinal, not measured — they encode direction and rank, not effect sizes. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Applied at Future Proof

How Future Proof™ applies this — gamify the evidence-backed action.

We take the meta-analytic caveats literally, so every mechanic is bolted to retrieval, never to a proxy for it. Streaks count consecutive days of actually recalling material — not logins, not video-completions — so the habit the streak protects is the habit that builds memory. XP is earned by answering retrieval questions and clearing mastery checks, which keeps the points pointed at learning outcomes rather than time-on-task. And team leaderboards are framed as committed, specific goals in the goal-setting sense the evidence supports, and scoped to teams to blunt the intrinsic-motivation risk that raw individual ranking carries. The mechanic amplifies; the retrieval does the learning.

See how the mechanics work
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Research Library PDF.

The evidence, by year

  • 1999Deci
  • 2014Hamari
  • 2017Mekler
  • 2017Sailer
  • 2017Landers
  • 2019Koivisto
  • 2020Sailer
  • 2020Bai
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1999–2020, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Hamari, J., Koivisto, J., & Sarsa, H. (2014). Does gamification work? — A literature review of empirical studies on gamification. Proceedings of the 47th Hawaii International Conference on System Sciences (HICSS): 3025–3034. DOIPDF
  2. Sailer, M., & Homner, L. (2020). The gamification of learning: a meta-analysis. Educational Psychology Review 32(1): 77–112. DOI
  3. Mekler, E.D., Brühlmann, F., Tuch, A.N., & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior 71: 525–534. DOIPDF
  4. Deci, E.L., Koestner, R., & Ryan, R.M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin 125(6): 627–668. DOI
  5. Bai, S., Hew, K.F., & Huang, B. (2020). Does gamification improve student learning outcome? Evidence from a meta-analysis and synthesis of qualitative data in educational contexts. Educational Research Review 30: 100322. DOI
  6. Sailer, M., Hense, J.U., Mayr, S.K., & Mandl, H. (2017). How gamification motivates: An experimental study of the effects of specific game design elements on psychological need satisfaction. Computers in Human Behavior 69: 371–380. DOI
  7. Landers, R.N., Bauer, K.N., & Callan, R.C. (2017). Gamification of task performance with leaderboards: A goal setting experiment. Computers in Human Behavior 71: 508–515. DOI
  8. Koivisto, J., & Hamari, J. (2019). The rise of motivational information systems: A review of gamification research. International Journal of Information Management 45: 191–210. DOI
See the mechanics

Gamify the action the evidence endorses — not a proxy for it.

Book a 20-minute demo using your team’s actual content. We’ll show you where the streaks, XP and leaderboards attach — always to retrieval practice — and why that keeps the game amplifying learning instead of gaming a metric.

8 citations Reviewed August 2026 Open peer review welcomed