Skip to content
The 21st Century Learning Initiative

How Accurate Are AI Detectors? What the Research Shows Teachers

A teacher opens the marking queue and finds a number beside a student's essay: 87 percent likely to be machine generated. The number arrives with the authority of software and none of the context needed to weigh it. What it means, how often it is wrong, and whom it is wrong about, the interface does not say.

A laboratory balance on a wooden bench, one pan holding a stack of student essays and the other a single brass weight, the beam tipped. Engraved duotone plate, ink blue on cream paper.
Plate XA laboratory balance on a wooden bench, one pan holding a stack of student essays and the other a single brass weight, the beam tipped. Plate drawn in the archive's ink and paper style.

This article is an evidence review. It asks how accurate AI detectors are in the sense that matters to a school: when a detector flags student writing, how much confidence does the flag deserve? No product is named; the findings hold across the category.

The short answer. Not accurate enough to decide anything on their own. Independent testing found no tool above 80 percent accuracy, error rates rose sharply against paraphrased text, and non-native English writers were flagged at high rates. A detector score can open a conversation. It cannot carry an accusation, and it cannot substitute for evidence of process.

What a detector actually measures

A text detector does not detect authorship. It measures statistical regularity. The two quantities most tools rely on are perplexity, how predictable each word is given the words before it according to a language model, and burstiness, how much that predictability varies across a passage. Machine text tends to be smooth: probable word choices, evenly built sentences, few surprises. Human text tends to lurch.

The inference the tool then makes is that smoothness implies a machine. That inference is the whole method, and it is weak, because smoothness has many causes. A writer with a limited vocabulary produces predictable text; so does a writer drilled in a formulaic genre, and so does a writer composing in a second language and choosing safe constructions. As this publication's essay on assessment when product severs from process argues, no inspection of the finished artifact restores the link to the person who made it.

How accurate are AI detectors in published testing?

The most thorough independent evaluation to date is Weber-Wulff and colleagues, 2023, in the International Journal for Educational Integrity. The team tested fourteen detection tools, including two commercial services licensed by universities, against human-written, machine-written, machine-translated, and paraphrased texts. None exceeded 80 percent overall accuracy, and only five exceeded 70 percent. The tools were also biased toward calling text human, so they miss much of what they are sold to catch while still raising false alarms.

The second finding concerns whom the errors fall on. Liang and colleagues, 2023, at Stanford, ran essays written by non-native English speakers for a standard proficiency test through seven widely used detectors. More than 60 percent were marked as machine generated on average, and 97 percent were flagged by at least one detector, while essays by native-speaking American schoolchildren were almost never flagged. When the researchers rewrote the non-native essays with more varied vocabulary, the flags mostly disappeared: the tools were reading fluency, not authorship. The companion piece on why detectors flag non-native English writers takes up the consequences.

The third data point came from the industry itself: in 2023 one major vendor of language models withdrew its own text classifier, citing low accuracy. If the company best placed to detect its own output cannot, a school should not assume a third party has.

Paraphrase evasion and the arms race

Even a detector that performed well on raw machine output faces a structural problem: the text it judges has usually been edited. Sadasivan and colleagues, 2023 showed that running machine text through a paraphrasing model, sometimes recursively, drove detection rates down sharply across several methods, in one case from about 70 percent to under 5 percent at a fixed 1 percent false positive rate. The paper also argued that as generated text comes to resemble human text more closely, the best achievable detector approaches a coin toss.

The upshot is an asymmetry. A student who wishes to evade detection can do so cheaply, by paraphrasing, translating, or editing by hand; an honest student cannot make their prose less predictable without writing worse. The detector falls hardest on those who did the work.

The arithmetic of a whole cohort

Vendors often quote a false positive rate of around 1 percent, and it is worth taking that figure at face value to see what it implies. A false positive rate is a rate per honest submission. Across a cohort it produces a stream of wrongly flagged honest students whose size depends on how many honest students there are, while the number of cheats caught depends on how many cheats there are and how many the tool can still see after paraphrasing. When honest students greatly outnumber dishonest ones, false alarms can outnumber true catches. The table runs the numbers for 500 submissions, assuming 70 percent sensitivity on unedited text (the better tools in Weber-Wulff) and a generous 30 percent on paraphrased text (after Sadasivan).

Machine-written shareHonest flagged, 1% FPRHonest flagged, 2% FPRCheats caught, uneditedCheats caught, paraphrased
2% (10 of 500)51073
5% (25 of 500)510188
10% (50 of 500)593515
20% (100 of 500)487030

Read across the rows. Where few students are cheating, or where those who cheat paraphrase, the honest students wrongly accused outnumber the dishonest students caught. Even in the most favourable row the flagged pile contains several innocent students, and the teacher holding one flagged paper cannot tell which pile it belongs to. And 1 percent is the advertised rate for native English writers; Liang's results put the rate for non-native writers far higher.

What a detector score is good for

None of this means a score carries no information; it means the information is weak and unevenly distributed, which fixes the role the score can play. A high score is a reasonable prompt for a conversation: a teacher may ask the student to talk through the piece, show the drafts, or explain a choice in the third paragraph. That takes minutes and usually settles the question.

What a score cannot be is evidence in a hearing. A panel that convicts on a percentage has accepted the testimony of a witness that cannot be cross-examined, whose error rate is intolerable, and whose errors fall disproportionately on students learning English. Our guidance on what a wrongly accused student can show exists because schools have reached findings this way. A policy that names the score as a trigger for inquiry and forbids it as grounds for a finding removes the risk; our working policy template includes that clause.

Process evidence is what settles it

The reliable evidence of authorship was never in the finished text. It is in the process: the notes, the rough first draft, the version history that shows a document growing over days rather than arriving in one paste, and the student's ability to defend the work aloud. This publication has argued since Collins, Brown and Holum's 1991 essay on cognitive apprenticeship that learning is made trustworthy by making thinking visible; the same principle answers the integrity question.

The tools for this are ordinary and mostly free. Tracked documents supply a version history that stands as evidence of authorship at no cost beyond a glance. A three-minute oral defence, sampled rather than universal, is difficult to fake and diagnostically rich. Each gives the student a way to demonstrate the work rather than a way to be suspected of it. The heaviest reliance on detectors falls on homework, the largest mass of unsupervised product a school assigns; the case for rethinking it is made in the homework question, asked honestly.

Frequently asked questions

What false positive rate do AI detectors have?

Vendors commonly advertise around 1 percent on native English prose; independent studies have measured far higher rates for particular groups, and Liang and colleagues found more than 60 percent of non-native writers' essays flagged. Across a whole cohort, even a 1 percent rate produces a steady stream of wrongly flagged honest students, often more than the cheats caught.

Can a student be disciplined on a detector score alone?

A school can, and some have, but the evidentiary basis is poor and increasingly challenged. A percentage from an unauditable tool with a known error rate does not meet the standard a reasonable panel should apply. Sound policy treats the score as grounds for a conversation and requires process evidence before any finding.

Why do AI detectors flag non-native English writers?

Because the tools measure predictability of word choice, and writers working in a second language tend to choose safe, common constructions. That produces low perplexity, the same signature the tools attribute to machines. Rewriting the same essays with more varied vocabulary removed most of the flags in the Stanford study, without changing who wrote them.

What should a teacher do with a high AI detector score?

Treat it as a reason to look, not a reason to decide. Ask the student to talk through the work, show the drafts or version history, and explain particular choices. That conversation usually settles the matter in minutes. Do not record the score as evidence, and do not frame the matter as an accusation before the student has had that chance to respond.

Where this leaves a school

The question of how accurate AI detectors are has a clear enough answer from the published research: not accurate enough to act on alone, unfair to some students, and easily evaded by anyone who intends to. A school may keep the tools as one prompt among several for a conversation with a student, and should write into policy that a score is never evidence in a hearing. The larger move is to stop asking the finished text to prove what it cannot, and to make the process visible while it happens. The Academic Integrity hub gathers the practical pieces of that redesign, and the essay on what academic integrity means when machines can write sets out the principles beneath it.

The human counterpart to this review, whether teachers can detect AI writing by eye, reaches the same verdict from the other side, and the essay on contract cheating after generated text explains why the evidence that once caught outsourcing has vanished with it.