© 2026 FUTURE PROOF™
AI & Tutoring · Randomized Evidence

LLM tutors: the first real RCTs.

For two years the argument about AI tutors ran on demos and anecdotes. Then the randomized trials arrived — from Harvard, the World Bank, and a thousand-student experiment in Turkey. What moved learning, what backfired, and why the tutor inside Future Proof™ is built to guide rather than answer.

TL;DR

The finding: The first randomized controlled trials of LLM tutoring show real, sometimes large gains. Harvard students learned more than twice as much with a purpose-built AI tutor as in an active-learning class. A six-week Nigerian pilot moved combined outcomes by roughly 0.3 standard deviations. But the same literature contains a warning: unrestricted ChatGPT access left students worse off on the unassisted exam.

The mechanism: When the model hands over finished answers, students substitute it for thinking, and the practice gains evaporate the moment the tool is removed. When it is constrained to scaffold — hints, questions, feedback on attempts — practice becomes effortful retrieval, and the gains survive.

The product: The tutor inside Future Proof is guardrailed to guide rather than answer — the design choice the RCT literature keeps validating.

In this article

  1. 01The bar tutoring has to clear
  2. 02Harvard: the AI tutor beat an active-learning class
  3. 03Nigeria: six weeks that moved the distribution
  4. 04Turkey: the trial that should worry everyone
  5. 05Reading the three trials as one experiment
  6. 06Design lessons the trials agree on
  7. 07What the evidence doesn’t show
© 2026 FUTURE PROOF™
The route. 7 sections, from “The bar tutoring has to clear” to “What the evidence doesn’t show”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Education technology has a long history of arriving before its evidence does. The pattern is always the same: a demonstration that astonishes, a wave of adoption that outruns measurement, and then — years later — trials whose results are quieter and stranger than either the enthusiasts or the skeptics predicted. Large language models compressed that cycle into about thirty months. This article covers what the compression produced: the first randomized trials of LLM tutoring, what they actually found, and the single design variable that keeps deciding whether the technology helps or harms.

For roughly two years after ChatGPT launched, the argument about AI tutors ran almost entirely on demos and intuition. One camp saw Bloom’s two-sigma dream finally within reach: a competent tutor for every learner on earth, at near-zero marginal cost. The other camp saw an answer machine that would do students’ homework for them and hollow out the very skills it claimed to teach. Both camps had screenshots. Neither had a control group.

That changed between 2024 and 2025, when the first genuine randomized controlled trials of LLM-based tutoring were run — in a Harvard physics course, in after-school classrooms in Nigeria, and in a Turkish high school. The results are more instructive than either camp predicted. The gains, where they appear, are large. The harm, where it appears, is real. And the variable that separates the two is not model quality. It is interaction design.

The bar tutoring has to clear

Tutoring is arguably the most-studied intervention in education. Bloom famously reported that students tutored one-to-one under mastery-learning conditions performed about two standard deviations above a conventional classroom (Bloom, 1984). That result was so striking it has framed the field’s ambitions for forty years. Yet the broader experimental record puts human tutoring well below it. A systematic review of 96 randomized evaluations of preK-12 tutoring programs found a pooled effect of roughly 0.37 standard deviations (Nickow, Oreopoulos & Quan, 2020) — far short of two sigma, yet still among the largest reliable effects in education.

Software tutors were closing in on that bar before LLMs existed. VanLehn’s review found intelligent tutoring systems produced learning effects nearly indistinguishable from human tutors — 0.76 versus 0.79 standard deviations in his synthesis (VanLehn, 2011). A later meta-analysis reported a median effect of 0.66 standard deviations across ITS evaluations (Kulik & Fletcher, 2016). So the open question for LLM tutors was never whether software can teach. It was whether a general-purpose conversational model — flexible, cheap, and unscripted — could match systems that took years to hand-engineer. And could it do so without the obvious failure mode: a chatbot that will cheerfully just tell you the answer?

Harvard: the AI tutor beat an active-learning class

The cleanest positive result so far comes from an introductory physics course at Harvard. Kestin and colleagues ran a randomized crossover (Kestin et al., 2025). In one week, half the students worked through a lesson with an AI tutor at home, while the other half covered the same material in a research-validated active-learning classroom. The following week, the groups swapped. This matters, because active learning — not passive lecture — is the strongest classroom baseline instruction has to offer.

The AI condition was not vanilla ChatGPT. It was a GPT-4-based tutor deliberately engineered around learning science: content broken into steps, guidance instead of immediate solutions, management of cognitive load, and growth-mindset feedback. Against the active-learning classroom, students in the AI-tutored condition showed markedly higher learning gains on the post-tests — the authors report more than double (Kestin et al., 2025). They reached those gains in less self-reported time on task, with higher engagement and motivation.

The detail that is easy to miss: the pedagogy was in the system, not the model. The authors are explicit that the tutor’s prompt encoded decades of physics-education research. The trial tests a designed tutor, not a raw chatbot — a distinction the next two studies make painfully concrete.

d = 0 0.5 1.0 1.5 2.0 Bloom’s 1:1 legend (1984) d ≈ 2.0human tutoring, 96 RCTs 0.37intelligent tutoring systems 0.76human tutors, same review 0.79ITS, later meta-analysis 0.66standard-deviation gains over conventional instruction © 2026 FUTURE PROOF™
Figure 1. The bar LLM tutors have to clear — the two-sigma legend that framed forty years of ambition, against what randomized evidence shows for human tutoring and for the pre-LLM intelligent tutoring systems that nearly matched it. Schematic after Bloom (1984), Nickow et al. (2020), VanLehn (2011) and Kulik & Fletcher (2016); read the contrast, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Students learn more than twice as much in less time when using an AI tutor, compared with the active learning class. Kestin et al. 2025, Scientific Reports

Nigeria: six weeks that moved the distribution

A single elite-university crossover, however clean, cannot carry a global claim — so the second question is whether anything similar happens where resources, connectivity, and baseline attainment all look completely different.

The Harvard result comes from one of the most-resourced classrooms on earth. The World Bank pilot in Edo State, Nigeria, sits at the other end of the spectrum. De Simone and colleagues randomized senior-secondary students into a six-week after-school program in which students used Microsoft Copilot (built on GPT-4) as an English tutor, with a teacher present to orchestrate the sessions (De Simone et al., 2025).

Treated students outperformed controls by roughly 0.3 standard deviations on a composite of English, digital skills, and AI knowledge. The authors benchmark that gain as comparable to one and a half to two years of business-as-usual schooling, and rank the program among the more cost-effective interventions in their comparison set (De Simone et al., 2025). Two secondary findings are worth flagging: effects grew with the number of sessions attended, and girls — who started behind boys — closed much of the gap.

The honest caveat is that the control group received no program at all, so the estimate bundles the AI tutor together with structured after-school time and teacher attention. The trial cannot isolate the model’s contribution. But as an existence proof that a guided GPT-4 deployment can move learning in a low-resource setting, in six weeks, it is hard to dismiss.

Turkey: the trial that should worry everyone

Two positive results in a row invite a comfortable conclusion — give students a capable model and learning follows. The third trial exists to demolish that conclusion, and it does so with the most consequential design choice in the early literature: it tested the same model twice, once with guardrails and once without.

The third anchor study is the reason “AI tutor” and “chatbot access” must never be used interchangeably. Bastani and colleagues ran an RCT with roughly a thousand high-school students in Turkey during math practice sessions (Bastani et al., 2024). Three arms: standard practice with notes and textbook; “GPT Base,” a stock ChatGPT-4 interface; and “GPT Tutor,” the same model wrapped in guardrails — it gave hints rather than answers, and carried teacher-designed prompts to reduce errors.

During practice, both AI arms looked terrific: performance improved by 48% with GPT Base and 127% with GPT Tutor relative to control. Then the tools were taken away for a closed-book exam. Students who had practiced with GPT Base scored 17% worse than students who had never touched AI at all. The guardrailed GPT Tutor group’s harm was essentially eliminated — they performed about the same as control — though notably, their large practice advantage did not convert into lasting gains either (Bastani et al., 2024).

The number

−17% Closed-book exam performance of students who had practiced with stock ChatGPT, relative to peers who never touched AI — the practice gains reversed the moment the tool left the room (Bastani et al., 2024).

The interaction logs explain the damage: students used the base model as an answer engine — paste the problem, copy the solution — and the model’s solutions were themselves frequently wrong. The authors’ framing is that students were using the tool as a crutch, and the crutch was load-bearing. Same model family as Harvard’s success story. What differed was what the system permitted.

Reading the three trials as one experiment

Set side by side, the three anchor studies form something the field did not plan but should notice: a natural gradient in interaction design. At one end, Harvard’s tutor was engineered top to bottom around learning science, and it beat the best classroom baseline available (Kestin et al., 2025). In the middle, Nigeria deployed a general-purpose assistant inside a structured, teacher-orchestrated session, and moved the distribution by a solid margin (De Simone et al., 2025). At the other end, Turkey’s unguarded arm handed students raw model access and left them worse off than no AI at all (Bastani et al., 2024). The guardrailed arm — same students, same model, same weeks — neutralized the harm.

The outcomes track the gradient, not the geography, the population, or the model family. That is about as close as early-stage literature gets to a causal statement about design. As pedagogical structure is removed, the effect decays from strongly positive through modest to negative. The sign flips exactly where the system stops withholding answers. It also explains why blanket verdicts on “AI in education” keep talking past each other — the category spans both ends of the gradient, and averaging across them produces a number that describes no deployment anyone actually runs.

Design lessons the trials agree on

1. The guardrail is the treatment.

Bastani’s two AI arms are, in effect, a randomized comparison of a single design variable: whether the model will hand over the answer. That one variable spanned the distance between the worst outcome in the study and a neutral one (Bastani et al., 2024). Kestin’s positive result sits on the same side of the line — its tutor was instructed to scaffold, not solve (Kestin et al., 2025). Across the early literature, “guide, don’t answer” is the closest thing to a confirmed design law.

Design rule

Guide, don’t answer. The same model, in the same classrooms, spanned the distance between the study’s worst outcome and a neutral one on a single design variable: whether the system would hand over the solution. Evaluate any AI tutor on what it refuses to do.

2. Pedagogy lives in the system, not the model.

None of the successful deployments was a raw chatbot. Consistent with this, Pardos and Bhandari found that LLM-generated hints, delivered inside the structure of a conventional tutoring system, produced learning gains comparable to hints authored by human tutors (Pardos & Bhandari, 2024). The model can supply the content; the scaffold decides whether the content teaches.

3. AI can raise the floor of human tutoring.

The tutor does not have to be the AI. Wang and colleagues randomized human tutors to receive Tutor CoPilot, a real-time assistant suggesting expert-like moves mid-session. Students whose tutors had it were about four percentage points more likely to master the session’s topic, and about nine points more likely when their tutor was lower-rated (Wang et al., 2024). The gains concentrate exactly where expertise is scarcest.

Why it matters

The assistant does not have to replace the tutor. Real-time AI coaching lifted mastery by about four percentage points overall and about nine when the human tutor was lower-rated (Wang et al., 2024) — the technology pays most where expertise is thinnest.

4. Access is not adoption, and adoption is not random.

In a massive online coding course, simply offering an LLM chat assistant reduced overall engagement, while the students who did adopt it went on to perform better on exams (Nie et al., 2024). Both halves matter: deployments cannot assume use, and observational comparisons of users versus non-users are hopelessly confounded by who chooses to adopt — which is precisely why this field needed randomized trials in the first place.

+150% +100% +50% 0 = control −50% +127% +48% GPT Tutor ≈0% GPT Base −17% during practice closed-book exam © 2026 FUTURE PROOF™
Figure 2. The Turkey reversal as a slope chart — with the assistant in hand both AI arms sit far above the no-AI control (+48% for the stock chatbot, +127% for the guardrailed tutor), but once the tools are removed for a closed-book exam the stock-chatbot line crosses below zero to −17% while the guardrailed tutor lands back at control. Data: Bastani et al. (2024); vertical axis is math performance relative to the no-AI arm. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Applied at Future Proof

How Future Proof™ applies this.

The AI Tutor runs under a hard guide-don’t-answer constraint — the same variable that separated the harm arm from the help arm in the Turkey trial. It never reveals a final answer. Instead it climbs a hint ladder (reframe the question → nudge the underlying concept → work one step, never the last), asks Socratic counter-questions, and requires the learner to commit an attempt before deeper help unlocks. Every assisted concept is flagged to the AI Memory Coach, which schedules an unaided retrieval of it days later — because the trials show assisted performance means nothing until it survives without the assistant.

See the AI Tutor

What the evidence doesn’t show

This literature is young, and it is worth being precise about its edges — six of them in particular.

  • No long-term retention data yet. The outcome measures in these trials sit days to weeks after the intervention. None measures what learners still know six months later — the outcome that actually matters — so we do not yet know whether LLM-tutored gains decay faster, slower, or the same as classroom-taught ones.
  • Nobody has shown two sigma. Harvard’s “more than twice the learning” is measured on lesson-level assessments in a one-week crossover, not a semester of standardized outcomes; Nigeria’s effect bundles extra instructional time. The pre-LLM ITS benchmarks of roughly 0.6–0.8 standard deviations (VanLehn, 2011)(Kulik & Fletcher, 2016) have not been formally surpassed at scale.
  • Effect sizes are not comparable across trials. The controls differ radically — an active-learning classroom, business-as-usual schooling, no program at all — so lining the numbers up side by side flatters some studies and punishes others.
  • Novelty effects can’t be excluded. Six weeks in Nigeria and one-week crossovers at Harvard are short windows; engagement with a new tool is at its peak in exactly that window.
  • Some of this is still working-paper science. The Turkey study circulated on SSRN, and several of the strongest adjacent results are preprints. Peer review and independent replication are in progress, not complete — in a field moving this fast, that caveat is not a formality.
  • The dependency question is open. The Turkey trial shows what one term of answer-on-demand does; nobody has run the multi-year trial that would tell us whether sustained guided-AI study makes learners better or worse at learning without it.

Where the evidence stops

  1. 1No long-term retention data yet
  2. 2Nobody has shown two sigma
  3. 3Effect sizes are not comparable across trials
  4. 4Novelty effects can’t be excluded
  5. 5Some of this is still working-paper science
  6. 6The dependency question is open
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The field’s first generation of RCTs, read together, does not say AI tutors work or fail. It says something more useful: they are a delivery mechanism whose sign — positive, null, or negative — is set by the pedagogy encoded around the model. That puts the burden exactly where it should be: on the design.

For a buyer evaluating an AI tutoring product, the practical translation is a short list of questions — questions the trials have already answered in the negative for naive deployments. Will it refuse to hand over the answer? Does it require an attempt before it helps? Is anything scheduled to check, later and unaided, whether the assisted learning survived? A vendor who cannot answer all three has built the arm of the Turkey trial that everyone should be trying not to build.

References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 1984Bloom
  • 2011VanLehn
  • 2016Kulik
  • 2020Nickow
  • 2024Bastani
  • 2024Pardos
  • 2024Wang
  • 2024Nie
  • 2025Kestin
  • 2025Simone
© 2026 FUTURE PROOF™
The evidence base. The 10 sources cited here span 1984–2025, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Bloom, B.S. (1984). The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring. Educational Researcher 13(6): 4–16. DOI
  2. Nickow, A., Oreopoulos, P., & Quan, V. (2020). The Impressive Effects of Tutoring on PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental Evidence. NBER Working Paper No. 27476. DOI
  3. VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist 46(4): 197–221. DOI
  4. Kulik, J.A., & Fletcher, J.D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research 86(1): 42–78. DOI
  5. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports 15. PDF
  6. De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. (2025). From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank Policy Research Working Paper No. 11125. PDF
  7. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics. SSRN Working Paper No. 4895486. DOI
  8. Pardos, Z.A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLoS ONE 19(5). PDF
  9. Wang, R.E., Ribeiro, A.T., Robinson, C.D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv preprint (2024). PDF
  10. Nie, A., Chandak, Y., Suzara, M., Malik, A., Woodrow, J., Sheeley, M., Piech, C., & Brunskill, E. (2024). The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters’ Exam Performances. arXiv preprint (2024). PDF
Try the AI engine

Meet the tutor that won’t hand over the answer.

Book a 20-minute demo with your team’s actual content. Bring a real problem, watch Future Proof’s tutor scaffold it — hints, counter-questions, attempt-first gating — and see the unaided follow-up it schedules to prove the learning stuck.

10 citations Reviewed August 2026 Open peer review welcomed