Can Teachers Detect AI Writing? What Reading Can and Cannot Show
Most teachers believe they would know. The essay arrives too smooth, the vocabulary sits a register above the student, and something in the rhythm of the paragraphs feels manufactured. The question this article asks is narrower: can teachers detect AI writing reliably enough, by reading alone, to act on what they sense?
The question matters because the feeling now carries consequences. A teacher who trusts her ear may open an integrity case on the strength of it, and a student who cannot prove a negative will carry the result. A teacher who distrusts her ear may grade a machine's fluency as a child's learning. Both errors come from treating an impression as a finding.
The short answer. Not by reading alone. In controlled studies, teachers separate machine-written from student-written essays at rates close to chance, and their confidence does not track their accuracy. What a reader notices is a reason to ask a question. Authorship is shown by the draft trail, the student's prior work, and a short conversation about the piece.
What experienced readers notice
Teachers are not imagining the signals. Anyone who has marked several hundred essays knows what fifteen-year-olds sound like, and generated prose departs from it in recognisable ways. The commonest departure is voice: a student whose supervised writing is halting and concrete submits a piece that is even, abstract and faintly corporate. The second is the absence of local error: a piece with no small faults, from a writer who usually produces them, stands out.
Other signals concern content. Generated essays are confidently general, asserting that "scholars have long debated" without naming a scholar. They may cite works that do not exist, because a language model produces plausible references rather than retrieving real ones. And they are often at odds with the room: a student who said nothing in Tuesday's discussion submits an essay that never mentions the source the class spent a week on. Each signal is real. None is proof.
Can teachers detect AI writing? What the studies found
The studies are consistent. Clark and colleagues, in a 2021 paper presented at the Association for Computational Linguistics, asked readers to distinguish human-written from machine-written stories, news articles and recipes. Untrained readers performed at roughly chance level, and brief training in what to look for improved accuracy only slightly. As the title All That's Human Is Not Gold suggests, readers took fluency and coherence, which a language model produces in abundance, as evidence of a human author.
The finding transfers to teachers. Fleckenstein and colleagues, in a 2024 study in Computers and Education: Artificial Intelligence, presented trainee and experienced secondary teachers with essays, some written by students and some by a language model. Both groups were unable to identify the generated essays reliably, performed close to chance, and were overconfident; experience did not help. Worse, teachers tended to rate the generated essays higher in quality, so the essays a teacher most wants to reward are the ones she is least able to attribute.
Software does not rescue the situation, as the review of how accurate AI detectors are shows, and the bias against non-native English writers that afflicts detectors afflicts human readers too: even register and generic phrasing are what a careful second-language writer produces on purpose.
Why the tells fail as evidence
Each tell has an innocent explanation at least as common as the guilty one. The student whose voice changed may have been coached by a parent, may have used a grammar checker, or may simply have tried harder on a piece that mattered. Fabricated citations are the most damning tell, and even they can be a student copying a reference badly from a secondary source. The table sets the common tells against how often, in an ordinary school, each would point to generated text rather than to something else.
| Tell | How reliable as evidence | Innocent explanation |
|---|---|---|
| Voice does not match the student's supervised writing | Moderate, where supervised work exists to compare | Coaching, a grammar tool, unusual effort |
| Even register with no local errors | Low: fluency is what proofreading produces | Proofreading, a careful second-language writer |
| Confident generalities without specifics | Low: the commonest fault of unaided student prose | Thin reading, weak technique, a vague prompt |
| Invented or unverifiable citations | High: hard to produce by accident | Miscopied reference, secondary source cited as primary |
| Essay ignores or contradicts the class discussion | Low to moderate | Absence, writing from the textbook |
| Formulaic structure: introduction, balanced points, summary | Very low | This is what students are taught to do |
A reader who sees several of these together has grounds to be curious, and nothing more. In a class where a handful of students used a machine and the rest did not, a tell present in half the generated essays but also in a fifth of the honest ones will point at more honest students than dishonest ones. That arithmetic is why a teacher's impression should start an inquiry rather than end one, and why the companion piece on what a student can show when falsely accused exists.
What actually shows authorship
If the product cannot carry the question, the process can. Three kinds of evidence do most of the work.
The first is the teacher's own knowledge of the student's writing over time. A single essay compared against nothing tells a reader little; the same essay compared against a term's worth of in-class writing and marked drafts tells her a great deal, and the comparison is visible and fair to the student. Schools that keep a process portfolio for each student have made the comparison routine.
The second is the draft trail. A document that grew over four evenings, with false starts and deleted paragraphs, looks unlike one that appeared in a single paste, and the difference is visible to anyone who opens the version history of a shared document. The lane's article on version history as evidence of authorship describes what the shapes look like and where they mislead.
The third is a conversation. Five minutes with the student and the essay, asking why the third paragraph was ordered that way, what was cut, what the second source argued, separates the writer from the copyist more decisively than any reading of the page. A student who wrote the piece talks about it as one talks about one's own decisions; a student who did not describes it from outside. The oral examination is the formal version; the informal version, sampled rather than universal, is the one most teachers can afford.
The teacher who watches the process does not need to guess
Collins, Brown and Holum's 1991 essay on cognitive apprenticeship and making thinking visible argued that the central difficulty of school learning is that thinking is hidden, and that instruction improves when it is brought into view. A master craftsman never needed to inspect a finished chair for signs that the apprentice had made it. He had watched it made and corrected the joints as they were cut. The judgment of authorship and the teaching were the same act.
Assessment that watches process recovers it; the site's statement on assessment when product severs from process sets out the design. A teacher who has seen the outline, read the first draft and talked to the student about the argument is not guessing from the product. The question "can teachers detect AI writing" only presses on a teacher who has arranged to see nothing but the finished text, and it is the arrangement that needs to change rather than the teacher's eye. As the older essay on the homework question noted before machines wrote anything, that is the arrangement most schools default to for most of the work they set.
Frequently asked questions
Can teachers tell if you used AI?
Not reliably, by reading alone. In controlled studies, teachers identify machine-written essays at rates close to chance and are overconfident in their judgments. What teachers can do is compare a submission with a student's supervised work, examine the draft history and ask the student about the piece. Those methods are far more reliable.
What are the signs of AI writing that teachers notice?
Experienced readers notice a voice that does not match the student's earlier work, prose with no local errors, confident generalisations without specifics, references that cannot be found, and essays that ignore what the class discussed. Each has an innocent explanation. Together they justify a question, never a verdict.
Should a teacher act on a strong feeling that an essay is AI-written?
Act on it by asking a question rather than by making an accusation. A strong feeling is a reason to look at the version history, compare the piece with earlier supervised work, and talk briefly with the student about the argument. If those checks support the impression, the school's policy applies. If they do not, the reader's ear was wrong, as it often is.
Is human judgment better than an AI detector?
Neither is good enough alone. Detectors report a probability without context, and human readers report an impression without calibration; both flag careful second-language writers and both reward fluency. A teacher's judgment becomes valuable when it is applied to drafts, prior work and conversation rather than to the finished page.
Where this leaves a school
Can teachers detect AI writing? Not from the page, and a school that expects them to has handed its teachers an impossible task and its students an unfair one. What teachers can do is watch the work happen, know their students' writing over time, and talk to them about what they made. Those practices establish authorship as a matter of record rather than suspicion, and they make detection largely unnecessary.
The Academic Integrity hub gathers the lane's articles on how that record is built. A teacher ready to begin should start with authentic assessment when the product can be generated, on setting work whose authorship answers itself.