Skip to content

The 21st Century Learning Initiative

Pillar / New Work, 2026

AI and Learning

This pillar asks the Initiative's founding question of a new subject: what does machine assistance change about how humans learn? Its position sits between hype and panic. AI tools neither end education nor improve it automatically; they move the scarce resource from finished answers to the effortful thinking a learner does before them.

6 briefings at launch · written for educators · every claim cited, every effect size named

Why This Pillar Exists

The Initiative's question, asked of AI

For thirty years the Initiative asked how humans learn, and pressed the answers on people who design schools. It never had to ask what happens when the finished products of thinking, the essay, the proof, the summary, can be produced without the thinking. That is now the ordinary condition of every classroom with an internet connection, and it deserves better than the two reflexes it usually gets: the vendor's promise that learning has been solved, and the columnist's verdict that learning has ended.

The briefings in this pillar hold to the archive's standard instead. Claims are cited. Effect sizes are named rather than gestured at. Predictions are written down plainly enough to be audited later, the way the Initiative's own 2010 predictions are audited elsewhere on this site. And every argument is measured against the research the archive already holds, because the questions machine assistance raises, about practice, visibility, and judgment, are old questions wearing new hardware.

The Archive's Lens

What cognitive apprenticeship predicts about AI tutoring

The most useful instrument this archive owns for judging AI tutors was published in 1991. Cognitive apprenticeship holds that expertise is learned by watching expert thinking made visible, then attempting the task with support that is deliberately withdrawn as competence grows. The withdrawal is not incidental. Fading is the mechanism; support that never fades is a crutch, whatever its interface.

Read that way, the 1991 model yields a working test that requires no benchmark suite. Does the tool show its reasoning, or only its answers? Does it ask the learner to articulate theirs? And above all, does it have any mechanism for stepping back? A system tuned to maximize helpfulness will fail the third question by design, which is why the model predicts that default settings, not model quality, will decide whether AI tutoring produces competence or dependence.

Plate V. Scaffolding fades; a crutch does not A line diagram with learner competence on the horizontal axis and support given on the vertical axis. One curve, labeled scaffolding, starts high and falls to zero as competence grows, which the 1991 model requires. A dashed flat line, labeled assistance that never fades, stays high across all levels of competence, which the model calls a crutch. PLATE V SCAFFOLDING FADES; A CRUTCH DOES NOT NOVICE COMPETENT THE LEARNER'S GROWING COMPETENCE SUPPORT GIVEN assistance that never fades the default setting of a tool tuned to be helpful scaffolding support withdrawn as competence grows Collins, Brown and Holum, 1991 the fade: the point of the whole design
Plate V. The 1991 model's design test for any tutor, human or machine: support must fall as competence rises.

The Evidence Base

The tutoring literature the AI industry now cites is worth reading at the source, because it is both stronger and stranger than the marketing suggests. Benjamin Bloom's famous 1984 claim put one to one tutoring two standard deviations above classroom instruction; later reviews could not reproduce an effect that size. Kurt VanLehn's careful 2011 synthesis found human tutoring nearer 0.79 and, remarkably, intelligent tutoring systems at 0.76, almost level. The honest reading is that structured tutoring works, machines can already deliver much of that structure, and nothing in the record says the benefit survives when the tutor does the thinking for the student.

Plate VI. What the tutoring research measured A scale of effect sizes in standard deviations from zero to two. Markers show 0.40 for human tutoring in the Cohen, Kulik and Kulik meta analysis of 1982, 0.76 for intelligent tutoring systems and 0.79 for human tutoring in VanLehn's 2011 synthesis, and 2.0 for Bloom's 1984 two sigma claim, marked hollow because later reviews could not reproduce it. PLATE VI WHAT THE TUTORING RESEARCH MEASURED 0 0.5 1.0 1.5 2.0 EFFECT SIZE, IN STANDARD DEVIATIONS ABOVE CLASSROOM INSTRUCTION 0.40 human tutoring, meta analysis Cohen, Kulik and Kulik, 1982 0.76 · 0.79 intelligent tutoring systems and human tutors, almost level VanLehn, 2011 2.0 the two sigma claim Bloom, 1984; not since reproduced the hollow marker is a claim; the solid markers are measurements
Plate VI. Tutoring effect sizes as the literature actually reports them, from the 1982 meta analysis to VanLehn's 2011 synthesis.

The Hard Question

Academic integrity and authentic student work

Nothing in this pillar teaches evasion, and nothing in it trusts detection. Those are the same position stated twice. Detection tools promise a certainty the underlying statistics cannot deliver, and their false positives fall hardest on honest students, disproportionately on those writing in a second language. A school that outsources its definition of integrity to a probability score has not solved its assessment problem; it has hidden it.

The archive suggests the older, sturdier route. Work is authentic when the thinking that produced it is visible: drafts that show revision, oral defence of a written argument, worked records of how a problem was attacked and abandoned and re attacked. Assessment built that way does not need a detector, because the process is the evidence. That is cognitive apprenticeship's principle pointed at grading, and the briefings below treat it as the design brief for the next decade of classroom practice.

The question is never whether a student used a machine. It is whether the work still contains the student.

From Proof of Work, briefing No. 03 in this pillar

The Pillar Index

New essays and briefings in this pillar

Six briefings at launch, written for educators, in the sourced and audited format the Initiative used for two decades. Each stands alone; together they are one argument about where learning lives when answers are cheap.

  1. The Apprentice and the Answer Machine

    What the 1991 model of modelling, coaching, and fading predicts about tutors that never tire and never step back.

    2026
  2. Writing to Think, When Machines Write First

    Drafting is where reasoning is built, not merely recorded; what remains of writing instruction when the first draft is delegated.

    2026
  3. Proof of Work: Assessment When Product Severs from Process

    When a finished product no longer proves its process, assessment has to watch the thinking: drafts, oral defence, worked records.

    2026
  4. The Homework Question, Asked Honestly

    The evidence for homework was thin before machine assistance; what independent practice is for, asked without nostalgia.

    2026
  5. AI Literacy for Educators: A Working Syllabus

    What these systems actually do, where they reliably fail, and how to read their output critically: a syllabus a department can run this term.

    2026
  6. What Tutoring Research Actually Says About AI Tutors

    Bloom's two sigma claim, the meta analyses that tempered it, and how the intelligent tutoring record reads before the sales pitch.

    2026

Where this pillar stands in the archive

These briefings are new work, but they are not a new question. They continue the inquiry the Initiative opened in 1996, and they answer to the same evidence base it spent three decades assembling.