AI Literacy for Educators: A Working Syllabus
AI literacy for teachers is the working knowledge needed to make sound professional judgments about machine assistance in classrooms: what large language models actually do, where they reliably fail, what learning science says about help and effort, and how to redesign assignments and assessment accordingly. It is a craft competence, not a technical specialty, taught here in six modules.
Why a syllabus and not a manifesto
Most of what teachers are offered about AI arrives as exhortation: adopt it all, or ban it all. Neither posture requires knowledge, which is why both are so easy to hold. The position of this publication, stated across the AI and Learning pillar, is that machine assistance in education is a question about how humans learn, and questions about how humans learn are answered by study, not by temperament. What follows is the syllabus we would set for a school staff taking that study seriously: six modules, each with a guiding question, core content, and an exercise, sized for a term of professional development meetings or an individual teacher's working through. It assumes no technical background. It does assume the reader teaches, and will test everything against their own classroom.
Module one: what the machine is doing
The guiding question: what kind of object is this? A large language model is a statistical system trained on enormous quantities of text to predict, given a passage, what plausibly comes next. Everything else follows from that. The system has no store of vetted facts, no model of the student across the table, and no access to whether its output is true; it has patterns, extraordinarily rich ones, in how humans have written. This explains the twin behaviors every teacher meets in the first week. The fluency is real, because fluent continuation is precisely what the system was trained to produce. The fabrication is equally real, because a plausible next sentence and a true one are different targets, and the system was trained on the first. Teachers should learn the word hallucination as a term of art for confident fabrication, and should notice that the machine's register never changes when it happens: the false citation arrives in the same assured prose as the sound summary.
The exercise for this module is direct experience. Ask an AI writing system ten questions inside your own subject expertise, where you can grade the answers, and score them. Then ask for sources and check them. Most teachers emerge with the correct calibration in one sitting: impressive competence, unreliable authority, and no way to tell one from the other by tone.
Module two: where the systems reliably fail
The guiding question: what should I never delegate? Fluency conceals a failure profile that is stable enough to teach. The systems fabricate specifics: citations, quotations, dates, numbers, and biographical details are unsafe without verification. They are weakest exactly where training text is thinnest, which includes recent events, local matters, and specialized corners of every discipline. They comply with a request's framing rather than its intent, so a flawed premise is usually elaborated, not corrected. Their arithmetic and multi-step logic can fail without warning signs. And their output regresses toward the generic, because prediction from the whole of written text is a machine for producing the average of it; the polished sameness teachers now recognize in submitted work is not a bug but the native style.
Each discipline should localize this map, because the failure profile lands differently by subject. In history, fabricated quotations and invented page references are the standing hazard. In mathematics, a correct method can be narrated around a wrong calculation so smoothly that the error hides in plain sight. In the sciences, the systems blend textbook consensus with confident overstatement at exactly the frontier where a syllabus wants precision. In literature, the machine produces readings of astonishing evenness, plausible about every text and urgent about none.
The classroom consequence is a rule of thumb worth posting in the staff room: the machine is most dangerous where you are least able to check it. A teacher fluent in their subject can use these systems productively because errors announce themselves; a student who has not yet learned the subject cannot, which is an argument, developed throughout this syllabus, for keeping the machine on the expert's side of the desk far more than the novice's.
Module three: the learning science baseline
The guiding question: what does the research say about help? This is the module that separates AI literacy for educators from AI literacy in general, and it is mostly not about machines. The relevant findings predate them. Retrieval practice strengthens memory more than restudying, the testing effect Roediger and Karpicke demonstrated in 2006. Conditions that make practice harder now often improve learning later, Robert Bjork's desirable difficulties, which warn that a tool minimizing a learner's effort may be minimizing the learning. Tutoring helps, but the celebrated two standard deviation claim from Bloom's 1984 paper did not survive review: VanLehn's 2011 synthesis put human tutors near 0.79 standard deviations and step-based tutoring systems near 0.76, a sobering and useful pair of numbers for anyone evaluating an AI tutor's marketing; the full argument is in What Tutoring Research Actually Says About AI Tutors.
The organizing framework we recommend is the one this publication was, in a sense, built around: the 1991 cognitive apprenticeship model of Collins, Brown, and Holum, in which learning proceeds by modelling, coaching, scaffolding, and fading. Read against that model, machine assistance sorts itself: valuable where it makes thinking visible, corrosive where it makes thinking unnecessary, and always in need of the one move it cannot perform, the deliberate withdrawal of support. The application of this framework to AI is worked through in The Apprentice and the Answer Machine, which pairs with this module.
Module four: assignment and assessment redesign
The guiding question: what still works? Machine text severed the finished product from the process that was supposed to produce it, which quietly broke the evidentiary basis of unsupervised graded work. The redesign principles are treated at length in Proof of Work and, for composition specifically, in Writing to Think, When Machines Write First; the syllabus version is brief. Make process the deliverable: drafts, notes, version history, and short oral defenses carry the evidence that products no longer can. Anchor writing in what only the student possesses, the class discussion, the local data, their own prior draft. Protect a core of unassisted, in-view practice, because module three says the effort is the mechanism. Apply the same honesty to homework, whose accounting function is gone even though its learning function survives, as argued in The Homework Question, Asked Honestly. The exercise: take your own most assignable task and redesign it twice, once to be machine-resistant, once to use the machine deliberately, and note which redesign taught you more about the task. Most teachers find the second redesign harder and more instructive, because deliberate use forces a question the first redesign lets you defer: what, exactly, is the irreducible thinking this task exists to produce? A teacher who can answer that question for every major assignment has done the deepest work this syllabus asks, and has usually discovered that a few cherished assignments cannot answer it, machine or no machine.
Module five: integrity without inquisition
The guiding question: what do I do about misuse? Two findings anchor this module. First, automated detection of machine writing is statistically unreliable in exactly the way that matters: a 2023 Stanford study led by Weixin Liang found detectors flagged most essays by non-native English speakers as machine generated, because they measure fluency patterns, not authorship. A false accusation rate like that is disqualifying for disciplinary use, and a teacher's intuition about voice, while diagnostically useful, is not proof either. Second, the productive posture is structural, not forensic: clear disclosure rules stating what assistance is permitted per task, process requirements that make authorship demonstrable rather than deniable, and conversations that treat a suspicious submission first as an assessment problem, what does this student actually know, which a five minute oral check answers better than any detector. Integrity policy written this way converts most incidents into teaching. Policy written as detection and punishment converts teaching into litigation, and loses.
Module six: judgment, policy, and the professional stance
The guiding question: what is my considered position? The final module asks teachers to write one page: where machine assistance belongs in their subject, where it does not, and why, with reasons drawn from the previous five modules rather than from mood. Staff who compare pages will disagree productively, and a department's collected pages are the honest first draft of its AI policy, better than anything adopted from a vendor's template or another school's panic. Teachers with policy influence should extend the exercise to the questions institutions must answer, the subject of our briefing Seven Questions for the Select Committee. The stance this publication models is the one its founders took toward every prior technology of learning: neither boosterism nor ban, but the steady application of what we know about how humans learn. That stance is itself teachable, and the position paper is how it gets taught. A teacher who has written one can defend a decision to a parent, a principal, or a student with reasons instead of rules, which is what professional judgment sounds like from the outside. A teacher who has not written one will borrow a position under pressure, usually the loudest one available that week.
How to run the term
As professional development, the syllabus fits a standard term: one module per cycle of meetings, with the exercises done between sessions and discussed in them, since the exercises, not the readings, are where calibration forms. The reading list is deliberately short and mostly old. The restored cognitive apprenticeship article anchors module three; VanLehn's 2011 review and Black and Wiliam's Inside the Black Box support modules three and four; the essays of this pillar supply the connective argument. A department can run the sequence alone, but mixed-subject groups are better, because module two's failure mapping and module six's position papers gain most from colleagues who check each other's blind spots. The end-of-term artifacts are concrete: a subject-specific failure map, two redesigned assignments per teacher, a draft disclosure policy, and the one page positions. A school that possesses those four documents has replaced anxiety with equipment.
Run individually, the same sequence takes a teacher perhaps twenty working hours across a term, most of them spent on the exercises. That is not a trivial ask, and it is worth stating why it earns its cost: every hour is spent inside the teacher's own subject and own assignments, so the output is not general awareness but usable judgment about Tuesday's lesson. Generic training in this field ages in months; calibration built on one's own discipline ages at the speed of the learning science underneath it, which is to say slowly.
Implications for practice
For an individual teacher, the syllabus compresses to a sequence of actions. Spend the hours of direct experimentation inside your own subject until your calibration is earned rather than borrowed. Learn the failure profile well enough to predict it. Put the testing effect, desirable difficulties, and the apprenticeship model in reach of your daily planning, because they, not the technology, decide what any tool is worth. Redesign your highest-stakes assignments so process carries the evidence. Replace detection with disclosure and oral assessment. Then write the one page position and revise it yearly, since the systems will change and the learning science largely will not. A teacher who has done these six things possesses AI literacy in the only sense that matters professionally: not the ability to operate the machine, but the ability to judge it.