LLM tutors: the first real RCTs.
For two years the argument about AI tutors ran on demos and anecdotes. Then the randomized trials arrived — from Harvard, the World Bank, and a thousand-student experiment in Turkey. What moved learning, what backfired, and why the tutor inside Future Proof™ is built to guide rather than answer.
The finding: The first randomized controlled trials of LLM tutoring show real, sometimes large gains — Harvard students learned more than twice as much with a purpose-built AI tutor as in an active-learning class, and a six-week Nigerian pilot moved combined outcomes by roughly 0.3 standard deviations. But the same literature contains a warning: unrestricted ChatGPT access left students worse off on the unassisted exam.
The mechanism: When the model hands over finished answers, students substitute it for thinking, and the practice gains evaporate the moment the tool is removed. When it is constrained to scaffold — hints, questions, feedback on attempts — practice becomes effortful retrieval, and the gains survive.
The product: The tutor inside Future Proof is guardrailed to guide rather than answer — the design choice the RCT literature keeps validating.
For roughly two years after ChatGPT launched, the argument about AI tutors ran almost entirely on demos and intuition. One camp saw Bloom’s two-sigma dream finally within reach: a competent tutor for every learner on earth, at near-zero marginal cost. The other camp saw an answer machine that would do students’ homework for them and hollow out the very skills it claimed to teach. Both camps had screenshots. Neither had a control group.
That changed between 2024 and 2025, when the first genuine randomized controlled trials of LLM-based tutoring were run — in a Harvard physics course, in after-school classrooms in Nigeria, and in a Turkish high school. The results are more instructive than either camp predicted. The gains, where they appear, are large. The harm, where it appears, is real. And the variable that separates the two is not model quality. It is interaction design.
The bar tutoring has to clear
Tutoring is arguably the most-studied intervention in education. Bloom famously reported that students tutored one-to-one under mastery-learning conditions performed about two standard deviations above a conventional classroom (Bloom, 1984) — a result so striking it has framed the field’s ambitions for forty years, even though the broader experimental record puts human tutoring well below that. A systematic review of 96 randomized evaluations of preK-12 tutoring programs found a pooled effect of roughly 0.37 standard deviations (Nickow, Oreopoulos & Quan, 2020) — far short of two sigma, yet still among the largest reliable effects in education.
Software tutors were closing in on that bar before LLMs existed. VanLehn’s review found intelligent tutoring systems produced learning effects nearly indistinguishable from human tutors — 0.76 versus 0.79 standard deviations in his synthesis (VanLehn, 2011) — and a later meta-analysis reported a median effect of 0.66 standard deviations across ITS evaluations (Kulik & Fletcher, 2016). So the open question for LLM tutors was never “can software teach?” It was whether a general-purpose conversational model — flexible, cheap, and unscripted — could match systems that took years to hand-engineer, without the obvious failure mode of a chatbot that will cheerfully just tell you the answer.
Harvard: the AI tutor beat an active-learning class
The cleanest positive result so far comes from an introductory physics course at Harvard. Kestin and colleagues ran a randomized crossover: in one week, half the students worked through a lesson with an AI tutor at home while the other half covered the same material in a research-validated active-learning classroom; the following week the groups swapped (Kestin et al., 2025). This matters, because active learning — not passive lecture — is the strongest classroom baseline instruction has to offer.
The AI condition was not vanilla ChatGPT. It was a GPT-4-based tutor deliberately engineered around learning science: content broken into steps, guidance instead of immediate solutions, management of cognitive load, and growth-mindset feedback. Against the active-learning classroom, students in the AI-tutored condition showed markedly higher learning gains on the post-tests — the authors report more than double — in less self-reported time on task, alongside higher engagement and motivation (Kestin et al., 2025).
The detail that is easy to miss: the pedagogy was in the system, not the model. The authors are explicit that the tutor’s prompt encoded decades of physics-education research. The trial tests a designed tutor, not a raw chatbot — a distinction the next two studies make painfully concrete.
Students learn more than twice as much in less time when using an AI tutor, compared with the active learning class.Kestin et al. 2025, Scientific Reports
Nigeria: six weeks that moved the distribution
The Harvard result comes from one of the most-resourced classrooms on earth. The World Bank pilot in Edo State, Nigeria, sits at the other end of the spectrum. De Simone and colleagues randomized senior-secondary students into a six-week after-school program in which students used Microsoft Copilot (built on GPT-4) as an English tutor, with a teacher present to orchestrate the sessions (De Simone et al., 2025).
Treated students outperformed controls by roughly 0.3 standard deviations on a composite of English, digital skills, and AI knowledge. The authors benchmark that gain as comparable to one and a half to two years of business-as-usual schooling, and rank the program among the more cost-effective interventions in their comparison set (De Simone et al., 2025). Two secondary findings are worth flagging: effects grew with the number of sessions attended, and girls — who started behind boys — closed much of the gap.
The honest caveat is that the control group received no program at all, so the estimate bundles the AI tutor together with structured after-school time and teacher attention. The trial cannot isolate the model’s contribution. But as an existence proof that a guided GPT-4 deployment can move learning in a low-resource setting, in six weeks, it is hard to dismiss.
Turkey: the trial that should worry everyone
The third anchor study is the reason “AI tutor” and “chatbot access” must never be used interchangeably. Bastani and colleagues ran an RCT with roughly a thousand high-school students in Turkey during math practice sessions (Bastani et al., 2024). Three arms: standard practice with notes and textbook; “GPT Base,” a stock ChatGPT-4 interface; and “GPT Tutor,” the same model wrapped in guardrails — it gave hints rather than answers, and carried teacher-designed prompts to reduce errors.
During practice, both AI arms looked terrific: performance improved by 48% with GPT Base and 127% with GPT Tutor relative to control. Then the tools were taken away for a closed-book exam. Students who had practiced with GPT Base scored 17% worse than students who had never touched AI at all. The guardrailed GPT Tutor group’s harm was essentially eliminated — they performed about the same as control — though notably, their large practice advantage did not convert into lasting gains either (Bastani et al., 2024).
The interaction logs explain the damage: students used the base model as an answer engine — paste the problem, copy the solution — and the model’s solutions were themselves frequently wrong. The authors’ framing is that students were using the tool as a crutch, and the crutch was load-bearing. Same model family as Harvard’s success story. What differed was what the system permitted.
Design lessons the trials agree on
1. The guardrail is the treatment.
Bastani’s two AI arms are, in effect, a randomized comparison of a single design variable: whether the model will hand over the answer. That one variable spanned the distance between the worst outcome in the study and a neutral one (Bastani et al., 2024). Kestin’s positive result sits on the same side of the line — its tutor was instructed to scaffold, not solve (Kestin et al., 2025). Across the early literature, “guide, don’t answer” is the closest thing to a confirmed design law.
2. Pedagogy lives in the system, not the model.
None of the successful deployments was a raw chatbot. Consistent with this, Pardos and Bhandari found that LLM-generated hints, delivered inside the structure of a conventional tutoring system, produced learning gains comparable to hints authored by human tutors (Pardos & Bhandari, 2024). The model can supply the content; the scaffold decides whether the content teaches.
3. AI can raise the floor of human tutoring.
The tutor does not have to be the AI. Wang and colleagues randomized human tutors to receive Tutor CoPilot, a real-time assistant suggesting expert-like moves mid-session: students whose tutors had it were about four percentage points more likely to master the session’s topic, and about nine points more likely when their tutor was lower-rated (Wang et al., 2024). The gains concentrate exactly where expertise is scarcest.
4. Access is not adoption, and adoption is not random.
In a massive online coding course, simply offering an LLM chat assistant reduced overall engagement, while the students who did adopt it went on to perform better on exams (Nie et al., 2024). Both halves matter: deployments cannot assume use, and observational comparisons of users versus non-users are hopelessly confounded by who chooses to adopt — which is precisely why this field needed randomized trials in the first place.
How Future Proof™ applies this.
The AI Tutor runs under a hard guide-don’t-answer constraint — the same variable that separated the harm arm from the help arm in the Turkey trial. It never reveals a final answer. Instead it climbs a hint ladder (reframe the question → nudge the underlying concept → work one step, never the last), asks Socratic counter-questions, and requires the learner to commit an attempt before deeper help unlocks. Every assisted concept is flagged to the AI Memory Coach, which schedules an unaided retrieval of it days later — because the trials show assisted performance means nothing until it survives without the assistant.
See the AI Tutor →What the evidence doesn’t show
This literature is young, and it is worth being precise about its edges.
- No long-term retention data yet. The outcome measures in these trials sit days to weeks after the intervention. None measures what learners still know six months later — the outcome that actually matters — so we do not yet know whether LLM-tutored gains decay faster, slower, or the same as classroom-taught ones.
- Nobody has shown two sigma. Harvard’s “more than twice the learning” is measured on lesson-level assessments in a one-week crossover, not a semester of standardized outcomes; Nigeria’s effect bundles extra instructional time. The pre-LLM ITS benchmarks of roughly 0.6–0.8 standard deviations (VanLehn, 2011)(Kulik & Fletcher, 2016) have not been formally surpassed at scale.
- Effect sizes are not comparable across trials. The controls differ radically — an active-learning classroom, business-as-usual schooling, no program at all — so lining the numbers up side by side flatters some studies and punishes others.
- Novelty effects can’t be excluded. Six weeks in Nigeria and one-week crossovers at Harvard are short windows; engagement with a new tool is at its peak in exactly that window.
- Some of this is still working-paper science. The Turkey study circulated on SSRN, and several of the strongest adjacent results are preprints. Peer review and independent replication are in progress, not complete — in a field moving this fast, that caveat is not a formality.
- The dependency question is open. The Turkey trial shows what one term of answer-on-demand does; nobody has run the multi-year trial that would tell us whether sustained guided-AI study makes learners better or worse at learning without it.
The field’s first generation of RCTs, read together, does not say AI tutors work or fail. It says something more useful: they are a delivery mechanism whose sign — positive, null, or negative — is set by the pedagogy encoded around the model. That puts the burden exactly where it should be: on the design.
Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
-
Bloom, B.S. (1984). The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring. Educational Researcher 13(6): 4–16. DOI
-
Nickow, A., Oreopoulos, P., & Quan, V. (2020). The Impressive Effects of Tutoring on PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental Evidence. NBER Working Paper No. 27476. DOI
-
VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist 46(4): 197–221. DOI
-
Kulik, J.A., & Fletcher, J.D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research 86(1): 42–78. DOI
-
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports 15. PDF
-
De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. (2025). From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank Policy Research Working Paper No. 11125. PDF
-
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics. SSRN Working Paper No. 4895486. DOI
-
Pardos, Z.A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLoS ONE 19(5). PDF
-
Wang, R.E., Ribeiro, A.T., Robinson, C.D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv preprint (2024). PDF
-
Nie, A., Chandak, Y., Suzara, M., Malik, A., Woodrow, J., Sheeley, M., Piech, C., & Brunskill, E. (2024). The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters’ Exam Performances. arXiv preprint (2024). PDF
Meet the tutor that won’t hand over the answer.
Book a 20-minute demo with your team’s actual content. Bring a real problem, watch Future Proof’s tutor scaffold it — hints, counter-questions, attempt-first gating — and see the unaided follow-up it schedules to prove the learning stuck.