Skip to content
The 21st Century Learning Initiative

Seven Questions for the Select Committee

Before regulating AI in schools, a legislature needs answers to seven questions: what problem the technology is solving, what the evidence predicts it will deliver, how supplier claims are verified, what happens to assessment, what data pupils surrender, what changes for teachers, and who is accountable when a system fails a child.

This briefing continues the Initiative's tradition of concise papers for parliamentarians, a tradition whose earlier entries, including the 2010 paper Schools in the Future, are preserved in the research archive. It is written for committee members preparing to take evidence on AI in schools policy. Each question is paired with the evidence a committee would need in hand to ask it well, and to recognize an inadequate answer.

Two remarks on method before the questions. First, the committee's difficulty will not be hostility from witnesses but fluency: the hearings will be full of confident, plausible testimony from parties with products to sell, and the questions below are designed so that a vague answer is visible as a vague answer. Second, the evidence base that matters most is not new. The strongest findings bearing on machine tutoring were published between 1982 and 2011, before the current technology existed, which is precisely what makes them credible: nobody who produced them had anything to sell. A committee that reads three papers, cited under question two, enters the room better armed than most of its witnesses expect.

1. What learning problem is this technology being asked to solve?

Witnesses will describe capabilities; the committee should require problems. Persistent gaps in early mathematics, unmet demand for individual tutoring, and teacher time consumed by administration are problems with evidence behind them. A capability in search of a problem is a procurement risk, not a policy. The committee should ask each witness to name the specific, measurable deficiency their proposal addresses, and what would count as failure.

2. What does the existing evidence predict it will deliver?

The relevant literature predates the current technology and is unusually clear. The famous claim that tutoring produces gains of two standard deviations comes from Benjamin Bloom's 1984 paper and has never been reproduced at scale. Meta-analytic estimates of real tutoring programs cluster far lower: 0.40 standard deviations in the classic 1982 synthesis by Cohen, Kulik, and Kulik, and 0.76 to 0.79 in VanLehn's 2011 comparison of intelligent tutoring systems with human tutors under favorable conditions. Our companion essay, What Tutoring Research Actually Says About AI Tutors, reviews this evidence in full. A committee that knows these numbers can calibrate every claim it hears: promised effects far above them are extraordinary claims requiring extraordinary evidence.

3. How will supplier claims be verified, and by whom?

Almost all current evidence for classroom AI products is supplier generated, short term, and measured on instruments aligned to the product. The committee should ask what independent evaluation is required before national deployment, whether evaluations use standardized measures and adequate comparison groups, and whether negative results must be published. A useful precedent is the requirement, common in medicines regulation, that trials be registered before results exist. Education has run this experiment before: reading schemes, interactive whiteboards, and one to one device programs all scaled nationally on supplier evidence, and the independent evaluations, where they came at all, came after the money. The committee is in a position to reverse that order this time, and should ask each witness directly whether they would accept independent evaluation as a condition of public purchase.

4. What happens to assessment when machines can produce the work?

Qualifications certify that a student did certain thinking. When fluent machine text is universally available, unsupervised written products no longer evidence that thinking, and detection tools cannot restore the link reliably. The committee should ask examination bodies how they will secure the connection between the credential and the candidate's own cognition, what mix of supervised, oral, and process based assessment they propose, and on what timetable, because the integrity of the qualification system is a national asset that erodes quietly.

5. What data do pupils surrender, and on what terms?

Adaptive systems necessarily record fine grained data about children's performance, errors, hesitations, and behavior. The committee should establish what is collected, where it is processed, how long it is retained, whether it trains commercial systems, whether it follows the child across providers, and what rights of deletion families hold. Children cannot consent meaningfully, schools are weak bargaining agents, and the data is uniquely sensitive: a record of a person's intellectual formation, gathered before they could object. The committee should also ask what happens to the record when a supplier is acquired, fails, or changes jurisdiction, since the lifetime of a company is shorter than the lifetime of the child it has profiled, and contracts silent on succession are answered later by liquidators.

6. What changes for teachers, and what must not?

The credible near term gains are in teacher workload: preparation, drafting, feedback, and administration. The risk is subtler, that judgment migrates to the system, and the teacher becomes an invigilator of machine decisions they cannot inspect. The committee should ask what training accompanies deployment, whether teachers can override and interrogate system recommendations, and how the profession's own expertise in the science of learning will be built rather than bypassed. The Initiative's long argument, from cognitive apprenticeship onward, is that learning depends on expert humans making thinking visible; policy should protect the conditions for that, whatever the tooling.

7. Who is accountable when the system is wrong about a child?

Every adaptive system misclassifies some learners: places them on the wrong track, predicts failure wrongly, flags honest work as machine made. The committee should ask, for each proposed use, who reviews adverse decisions, what appeal a family has, whether a human with authority can reverse the system, and where liability rests among supplier, school, and state. If witnesses cannot answer for the individual wrongly judged child, the deployment is not ready, whatever its average effects.

What a committee should do with the answers

Three dispositions follow from the evidence. First, calibrate: fund and permit deployments in proportion to the tutoring literature's realistic effect sizes, not the legend, and require independent evaluation as a condition of scale. Second, sequence: settle assessment integrity and data protection before mass procurement, because both are cheaper to design than to retrofit. Third, preserve the mechanism: whatever is bought, the policy test is whether pupils still do the cognitive work, since the entire evidence base for learning agrees on that point. Committees have been promised transformation by every educational technology since the teaching machines of the 1960s. The record suggests that the legislatures which served their schools best were those that asked short questions, required numbers, and kept the burden of proof where it belongs. Further briefings in this series appear in Briefings, alongside the new research essays in AI and Learning.