Educational AI

AI that tutors, adapts practice and grades work — and what forty years of effect-size research actually shows about whether it improves learning.

Published Updated

On this page

Definition

Educational AI — the phrase most people type is "AI in education" — is software that takes over part of a teacher's job: deciding what a student works on next, responding to what they just wrote, and judging whether they have understood it. The benchmark it has been chasing for forty years is Benjamin Bloom's 1984 report that one-to-one human tutoring moved the average student two standard deviations above a conventional classroom, from the 50th percentile to roughly the 98th; no computer tutor has ever come close, and the best-measured effects sit between 0.3 and 0.8 standard deviations.

That gap is the whole subject. Everything the category contains — intelligent tutoring systems, adaptive practice engines, automated essay scoring, early-warning analytics that flag a student at risk of failing, and the teacher-facing planning tools covered in the AI for lesson plans guide — is an attempt to buy some fraction of a private tutor at software prices. The interesting question is not whether these systems can produce a response that looks like teaching. Since large language models arrived, they obviously can. The question is whether the student ends up able to do the thing without the machine, and there the evidence points in both directions depending on one design decision.

The decision is what the system refuses to do. The same GPT-4, given two different sets of instructions, produced a measurable gain in one classroom trial and a measurable loss in another arm of the very same trial. If you take one thing from this page, take that: in educational AI the model is not the variable, the scaffold is.

How It Works

Two quite different architectures both get called educational AI, and they fail in different ways.

The classic intelligent tutoring system (ITS) is built around an explicit model of the student. Kurt VanLehn's 2006 description of these systems splits them into two loops. The outer loop runs once per task: it consults the student model — a running estimate of which sub-skills the learner has and has not mastered — and picks the next problem. The inner loop runs once per step within a problem: it checks the student's move, offers a hint, and updates the model. Carnegie Learning's Cognitive Tutor and ALEKS are the long-lived commercial examples.

The loop distinction is not academic bookkeeping, because granularity turns out to be where the effect lives — and not in the direction you would guess. VanLehn's 2011 meta-analysis of tutoring studies from 1975 to 2010 found human tutors raised scores by an effect size of 0.79 over no tutoring. Step-based computer tutors, which give feedback at the level of a solution step, reached 0.76 — statistically indistinguishable from a human. Substep-based tutors, which break the work down further still and comment inside each step, managed only 0.40. More granular help was worse. And none of them reproduced Bloom's 2.0.

The LLM tutor has none of this machinery. There is no student model, no problem bank, no skill graph. There is a general-purpose model, a system prompt, and a conversation. All of the pedagogy is compressed into instructions about what the model must not say — and because a model tuned for helpfulness will answer the question it was asked, those instructions are fighting the model's default behaviour rather than expressing it.

The cleanest measurement of what that is worth comes from a field experiment run in a large Turkish high school in autumn 2023 and published in PNAS in 2025 (Bastani, Bastani, Sungu et al.). Roughly 1,000 students across about 50 classes in grades 9–11 did four 90-minute maths practice sessions in one of three conditions: no AI, a plain ChatGPT-style interface ("GPT Base"), and a prompt-engineered tutor instructed to withhold answers and ask questions back ("GPT Tutor").

During practice, with the AI in front of them, control students averaged 0.284 out of 1.0. GPT Base scored 0.421 — a 48% improvement. GPT Tutor scored 0.645 — a 127% improvement. Then the researchers took the AI away and gave everyone an exam. GPT Base scored 17% below control (−0.054 out of 1.0, p < 0.05). GPT Tutor came in at −0.004: no harm, and no benefit either.

Read those four numbers together, because the headline usually reports only the first half. Access to an unguarded chatbot made students look more than twice as capable during practice and left them measurably worse off. The guardrail repaired the damage; it did not create a gain. The most confident-looking session in the room produced the weakest exam.

The mechanism has a name that predates chatbots by fifteen years. Koedinger and Aleven called it the assistance dilemma in 2007: giving help reduces frustration and saves time but can produce shallow learning, while withholding it forces the effortful retrieval that makes knowledge stick — at the cost of frustration and wasted minutes. Every tutoring system is a position on that trade-off. An LLM trained by reinforcement learning from human feedback to be maximally helpful sits, by default, at the end of the dial that teaches least.

Real-World Applications

Khanmigo is the largest schools deployment. Khan Academy reported it reaching 795 US districts and about 770,000 students by the end of the 2024–25 school year, with more than 108 million interactions since its 2023 launch. Its own published numbers are also the most useful corrective on the market: in May 2026 Khan Academy disclosed that only about 15% of students with access actually use it. A separate six-month internal test programme, run across more than 15 million tutoring threads between October 2025 and April 2026, found its biggest single win came from feeding the tutor a student's recent performance history — worth +6.1% on next-item correctness. That is a real improvement, and it is the scale of improvement the best-resourced team in the field found after twenty controlled tests: single-digit percentages on an in-product metric, not sigmas on an exam. As of mid-2026 no peer-reviewed randomised trial of Khanmigo's effect on learning has been published; the one independent comparison, 69 undergraduates studying lunar phases, found significant gains in every condition and no significant difference between Khanmigo, Google search, and paper.

The World Bank's Nigeria pilot is the strongest positive result from a developing-country setting: about 800 senior secondary students across nine public schools in Benin City, six weeks of after-school English sessions using Microsoft Copilot with a teacher present, at $48 per student, producing a 0.31 standard deviation gain. Read the comparison condition before you read the effect: the control group received no after-school programme at all, so the measured gain includes the extra instructional hours and the teacher supervising them.

The Harvard physics trial (Kestin et al., Scientific Reports, 2025) is the strongest result from a rich setting, and the most carefully bounded. In a crossover design, 194 undergraduates in Physical Sciences 2 alternated between an expert-run active-learning class and a purpose-built AI tutor at home. The AI condition produced more than double the learning gain in less time — a median of 49 minutes versus 60 — with effects reported in the range of 0.73 to 1.3 standard deviations. The authors' own caveats matter as much as the number: one topic, two weeks, Harvard undergraduates, no measure of retention or transfer, and a tutor designed by the same physicists who designed the class it beat.

Duolingo Max is the mass-market case, and a lesson in reading what was measured. Its 2025 Frontiers in Education study of 385 learners using the Roleplay and Explain My Answer features found significant gains in self-efficacy — learners' confidence that they could hold a real conversation. The authors state plainly that no measure of language-learning outcomes was included. Confidence is worth something in language learning, but it is not proficiency, and the study does not claim it is.

Key Concepts

An effect size is the unit this whole field argues in, and misreading it is how a school ends up buying the wrong thing.

  • A standard deviation is a distance, not a percentage. Bloom's 2σ means the average tutored student outscored 98% of the untutored class. A 0.3σ gain — the Nigeria result — moves the average student from the 50th percentile to about the 62nd. Both are worth having. They are not the same claim, and vendors quoting "two sigma" as an aspiration are quoting a number nobody has reproduced with software.
  • The counterfactual is the claim. "0.31 SD" is meaningless until you know what the other group did. Against nothing, most structured programmes look good. Against a textbook, a worksheet or an equivalent hour with a teacher, far fewer do. Ask what the control condition was before you ask what the effect was.
  • Performance during practice is not learning. The Turkish trial separates them cleanly: +48% while assisted, −17% once unassisted. Any metric a tool reports about activity inside itself — problems completed, time on task, accuracy in-session — measures the assisted number. Only an unassisted assessment measures the other one.

Challenges

The evidence base is thinner than its citation count suggests. The most-cited quantitative claim in this field was a 2025 meta-analysis in Humanities and Social Sciences Communications reporting a large effect of ChatGPT on learning performance (g = 0.867) across 51 studies. It accumulated hundreds of citations, then was retracted in 2026 after Magnus Ingebrigtsen and Marko Lukic documented miscounted studies, mis-weighted effects, a retracted study inside the sample, and inadequate control for confounds; their reading of the forest plot put the true figure nearer 0.5. Separately, 33 of the 51 included studies had fewer than 35 students in the ChatGPT group. Anyone who cited that 0.867 between May 2025 and the retraction was citing a number that did not survive scrutiny — and a great deal of procurement material still does.

Bloom's own 2σ is shakier than its fame implies. It rests on two 1984 University of Chicago dissertations by Joanne Anania and Joseph Burke; Anania's study had about 40 students, the tutors were undergraduate education majors trained for a week, and only one of the six experiments was genuinely one-to-one — the rest used one tutor to three students. It has never been replicated at scale. The target the industry has been chasing for forty years may simply be too high.

Access is not use. A tool that 15% of students open cannot produce a class-wide effect no matter how good it is when opened. Adoption, not capability, is the binding constraint in most district deployments, and it is the number least often reported.

Nothing measures retention. Almost every study above runs for weeks. The Turkish trial ran four sessions; the Harvard trial, two weeks; the Nigeria pilot, six. A tool that helps this term and leaves nothing behind next year would be invisible to all of them.

A hallucinated explanation is more damaging in a classroom than in an office, because the reader is by definition unable to check it — that is why they asked. And student data carries obligations that ordinary software does not; see Privacy and the FERPA constraints in the lesson-plans guide before putting a real student's name into any consumer AI account.

Randomised trials of the flagship products. The most consequential near-term result is a rigorous trial of Khanmigo, the first test of whether the largest deployed AI tutor changes outcomes at all. A null result from it would be more informative than any of the positive findings above, because it would be measured on a product millions of students already have.

Scaffold design becomes the research variable. The Turkish result reframed the question from "does AI tutoring work" to "what must the tutor refuse." Expect the next generation of studies to compare prompt and interface designs against each other rather than against no-AI controls — which is also the comparison a school actually faces, since the students already have the chatbot.

Assessment moves back into the room. If unsupervised work no longer distinguishes what a student can do from what their assistant can do, the assessment has to change rather than the policy. Oral defences, in-class writing and process evidence are the practical responses, and they cost teacher time that AI's efficiency gains are meant to be freeing up.

A higher bar for claims. After a widely cited meta-analysis was withdrawn, preregistration, active control groups and delayed unassisted post-tests are becoming the price of a credible efficacy claim in this field. That is the most useful thing that could happen to educational AI, and it will make the reported effect sizes smaller.

Frequently Asked Questions

Sometimes, and by much less than the marketing implies. The best-measured gains sit between 0.3 and 0.8 standard deviations, against the 2.0 that Benjamin Bloom reported for human one-to-one tutoring in 1984. The largest randomised trial of a general chatbot in schools — roughly 1,000 Turkish high-schoolers — found students who practised with an unrestricted GPT-4 scored 17% lower than the control group on a later exam taken without it.
Older adaptive software chose the next problem from a fixed bank using rules written by its designers. An intelligent tutoring system adds a student model that estimates which sub-skills a learner has mastered. A large language model tutor has neither: it generates its response from a general model, and its entire pedagogy lives in the instructions it was given about what not to say.
Because effort is the thing that produces learning, and a helpful assistant removes it. Learning scientists call this the assistance dilemma: help saves time and frustration but can leave the student unable to do the step alone. In the Turkish trial, students using the unrestricted chatbot solved 48% more practice problems and then lost ground on the unassisted exam — they had watched the work rather than done it.
Less than you would expect. As of mid-2026 no peer-reviewed randomised trial has shown that Khanmigo improves learning outcomes; the one independent comparison, with 69 undergraduates, found no significant difference between Khanmigo and using Google. Duolingo's published generative-AI study measured learner confidence, not language proficiency, and says so.
What the tool refuses to do, and what the control group in their evidence was doing instead. A gain measured against students who received nothing tells you the programme beats nothing; it does not tell you it beats a worksheet, a textbook or a teacher's time. Ask for the comparison condition before you ask for the effect size.
Nothing in the measured evidence supports it. Even the strongest result — a Harvard physics trial where an AI tutor produced more than twice the learning gain of an expert-run active-learning class — used a tutor built by the same instructors on the same lesson design, and covered a single topic over two weeks. The tools that show gains are ones a teacher configured, deployed and supervised.

Continue Learning

Explore our use-case guides and prompts to deepen your AI knowledge.