Skip to content
The 21st Century Learning Initiative

What Tutoring Research Actually Says About AI Tutors

Fifty years of tutoring research supports a modest, reliable claim: tutoring raises achievement by roughly 0.3 to 0.8 standard deviations, not the two standard deviations of legend. That is also what the evidence predicts about AI tutoring effectiveness: real gains, hard won, and dependent on the learner doing the thinking.

Every generation of educational technology has claimed the mantle of the personal tutor, and the current generation claims it more confidently than any before. The claim deserves to be tested against the literature that made tutoring famous in the first place. That literature is older, stranger, and considerably more sobering than the way it is usually cited.

Where the two sigma legend began

The citation behind almost every promise of transformative machine tutoring is Benjamin Bloom's 1984 paper in Educational Researcher, "The 2 Sigma Problem." Bloom reported that students taught one to one by a good tutor, under mastery learning conditions, performed about two standard deviations above students taught in a conventional classroom of thirty. Two standard deviations is an enormous figure: it would move the average student to roughly the 98th percentile of the untutored group.

The finding came from doctoral work conducted under Bloom's supervision at the University of Chicago, principally the dissertations of Joanne Anania and Arthur Burke. The studies were small and carefully controlled. They ran for a few weeks, in a narrow band of subject matter, with instruction and testing tightly aligned, and they combined tutoring with mastery learning, a method with its own independent effect. Bloom himself was scrupulous about what he was claiming. He presented two sigma as a problem, not a product: since one to one tutoring is too expensive to provide for every student, the research task was to find group methods that approach its results. The paper is a challenge to instructional design, and it has been quietly converted, in the decades since, into a benchmark that the original studies never established for the world outside their own conditions.

A result that did not reproduce

The two sigma effect has not been reproduced at scale. This is not a scandal; it is the ordinary fate of large effects found in small, tightly controlled studies. When tutoring is delivered across whole schools, by ordinary tutors, measured on tests that were not written to match the instruction, the effects shrink to a fraction of Bloom's figure. Reviews of well powered randomized trials in education find median effects nearer one tenth of a standard deviation, which is why an intervention that reliably delivers 0.3 or 0.4 counts as a major success in this field. Against that background, two full standard deviations was always an outlier ceiling, achievable in the laboratory sense that a record wind assisted sprint is achievable, and not a baseline any real program should be judged against or sold upon.

The honest summary is this: Bloom identified the upper bound of what individualized instruction might accomplish under ideal conditions. Everything since has been an attempt to find out how much of that upper bound survives contact with actual schools.

What the meta-analyses actually found

The most cited synthesis of real school tutoring programs came two years before Bloom's paper. Cohen, Kulik, and Kulik's 1982 meta-analysis in the American Educational Research Journal reviewed 65 evaluations of school tutoring programs and found an average effect on tutored students of 0.40 standard deviations. They also found something the two sigma story never mentions: the tutors themselves learned from tutoring, gaining about a third of a standard deviation in the subject they taught, and both groups came away with better attitudes toward the material. Effects were larger in structured programs, in shorter programs, and in mathematics rather than reading.

Later syntheses have landed in the same territory. A 2020 review by Nickow, Oreopoulos, and Quan, covering randomized field experiments of preschool through secondary tutoring, found a pooled effect of about 0.37 standard deviations, with the strongest results for tutoring delivered by teachers and trained paraprofessionals, embedded in the school day, in the early grades. Unstructured volunteer and parent tutoring performed markedly worse. The pattern across four decades is remarkably stable: tutoring works, it works to the tune of roughly a third to two fifths of a standard deviation in ordinary conditions, and its effectiveness depends heavily on structure, training, and integration with the curriculum.

VanLehn's correction

The most important paper for anyone reasoning about machine tutors is Kurt VanLehn's 2011 review in Educational Psychologist, which compared human tutoring, intelligent tutoring systems, and other computer tutoring against classroom instruction. Two findings matter here.

First, VanLehn found that human tutoring, measured across the accumulated comparisons, produced an effect of about 0.79 standard deviations, not 2.0. The legend had roughly doubled the reality even for human beings.

Second, the best step based tutoring systems of that era, systems that tracked a learner's progress through each step of a problem and gave feedback at that grain, produced an effect of about 0.76, statistically indistinguishable from the human figure. Software with no language ability and no charm matched human tutors on measured learning. VanLehn's explanation, the interaction granularity hypothesis, was that the benefit of tutoring comes largely from the fineness of the interaction, the tutor engaging with each step of the learner's reasoning rather than only with the final answer, and that beyond a certain granularity the returns flatten out. Human warmth, in the measured outcomes, added surprisingly little on top.

Why tutoring works when it works

The mechanisms behind these numbers are well described and worth stating, because they are the checklist against which any AI tutor should be examined. Tutoring provides immediate feedback at the point of error. It adapts the next task to the learner's current state. It elicits explanation: good tutors ask learners to say why, and the act of self explanation is itself a powerful driver of learning. And it holds attention, keeping the learner working at the edge of competence for more minutes per hour than a classroom can.

Readers of this archive will recognize the pattern. It is the cognitive apprenticeship cycle of modelling, coaching, scaffolding, and fading: the tutor makes thinking visible, supports the learner's own attempt, and then, critically, withdraws so the learner carries the reasoning alone. The withdrawal is not incidental. A tutor who never fades has not tutored; a tutor who supplies the answer has done something worse than nothing, because the learner has now practiced receiving conclusions rather than producing them. This distinction, between a system that elicits thinking and a system that dispenses answers, is the single most important variable the tutoring literature offers for the AI era, and it is the one that vendor demonstrations are least likely to exhibit.

What this predicts for AI tutors

Read plainly, the literature makes four predictions about AI tutoring effectiveness.

First, the realistic ceiling is the human tutoring figure, somewhere around three quarters of a standard deviation, achieved under favorable conditions: structured content, protected time, learners who show up, and a system that works at fine interaction granularity. Claims that exceed the ceiling of human tutoring should be treated as claims that the two sigma legend is true after all, and should carry the burden of proof that implies.

Second, most deployments will land well below the ceiling, in the 0.2 to 0.4 range that field studies of tutoring reliably produce, and some will produce nothing measurable, exactly as unstructured human tutoring produces little. The delivery conditions, scheduling, curriculum alignment, adult monitoring, will predict outcomes better than the sophistication of the underlying model.

Third, effects will be largest for learners who currently receive the least individual attention, and smallest for learners who already have skilled adults engaging with each step of their thinking. This is the equity case for machine tutoring, and it is a genuine one, provided access does not sort the same way every prior technology has sorted.

Fourth, an AI tutor that answers rather than asks will show inflated short term performance and depressed long term learning, because it removes the very effort that the tutoring mechanisms depend on. The literature on this point predates the machines and is unambiguous: assistance that substitutes for the learner's processing does not transfer.

What the literature does not license

It is worth being equally clear about the questions this evidence cannot answer. The tutoring studies measured achievement on tests administered soon after instruction; they say much less about retention over years, about transfer to unfamiliar problems, or about the slow formation of intellectual habits, which are the outcomes schooling actually exists to produce. The comparisons in VanLehn's review were dominated by well structured domains, mathematics, physics, and computing, where a problem decomposes into checkable steps. Whether the same effects extend to interpreting a poem, weighing historical evidence, or constructing an argument, domains where the steps are contested and the feedback is judgment, is not established by any of these numbers.

Nor does the literature speak to what a tutor does besides teach. Human tutoring in real programs has often carried effects that never appear in the effect size: an adult who notices a child weekly, expects things of them, and vouches for them. The 1982 meta-analysis's finding that tutors themselves learned and grew in confidence hints at this wider ecology. A machine that matches the instructional interaction replaces none of it. Schools deciding what to buy should hold both facts at once: the measured teaching effect may transfer to machines, and the unmeasured relationship certainly does not.

Finally, short studies flatter novelty. Effects measured in the first enthusiastic months of any technology have historically shrunk as the novelty wore off, and no long horizon studies of the current generation of tutors yet exist.

Implications for practice

For a school weighing an AI tutoring system, the research suggests a short list of questions. Ask what comparison group and what test sit behind any claimed effect, and discount alignment between the instruction and the measure. Expect a third of a standard deviation in good conditions and plan the conditions rather than trusting the tool: scheduled time, curricular integration, an adult who reviews progress. Insist on observing whether the system asks before it tells, and whether its scaffolding fades. And measure locally, because the honest range of outcomes runs from nothing to substantial, with implementation deciding where any given school lands.

None of this diminishes what may be possible. A technology that reliably delivered even half of VanLehn's human tutoring figure to every student who lacks a tutor would be among the most consequential interventions in the history of schooling. But the tutoring literature earned its authority through measured claims, and it extends that authority only to measured claims. The essays in the AI and Learning pillar continue this line of inquiry, and our briefings address what the same evidence means for the policymakers who will be asked to fund it.