Why AI Detectors Flag Non-Native English Writers
A student who learned English at fourteen writes a careful, correct, slightly stiff essay on the causes of the First World War. A detection service returns "likely AI generated". She has no draft to show, because she wrote it in one sitting the way she was taught, and the teacher has a screen that says 87 per cent. An ai detector false positive of this kind follows from what the detectors measure, and no software update will remove it.
These tools detect predictability, and anyone whose English is more predictable than a fluent native writer's, which describes most people writing in a second language, is more likely to be flagged. The consequence for international students and multilingual learners is structural, and the research published since 2023 has measured it.
The short answer. AI detectors estimate how easily a language model could have predicted each word of a text. Second-language writers use smaller vocabularies and more formulaic sentence shapes, so their prose scores as machine-like. Liang and colleagues (2023) found most non-native essays in their sample were flagged while native essays passed. Detection is a weak witness; process evidence is the reliable one.
What the detectors actually measure
The common design behind AI detectors runs the text through a language model of its own and asks, word by word, how surprised the model was. The average surprise is called perplexity. A text with low perplexity is one the model could have written itself, because each word sat near the top of its list of likely continuations. A text with high perplexity contains choices the model did not expect: an unusual verb, a regional idiom, a joke.
A second measure is usually layered on top. Burstiness describes how much the surprise varies across a passage. Human writing, on the detectors' theory, alternates between plain stretches and unexpected ones, while machine text sits at an even level of predictability throughout. Low perplexity and low burstiness together produce the verdict "likely AI", and the theory has some footing, since language models are built to emit the probable word.
Nothing in this method touches authorship. It knows only that the words were predictable, and it treats predictability as a proxy for machine origin. Any human whose prose is predictable for other reasons falls on the wrong side of that proxy, and the clearest such group is people writing in a language they are still acquiring.
Why second-language writing looks predictable
Vocabulary is the first constraint. A learner with a working vocabulary of a few thousand words chooses from a far smaller set than a native writer, and the words chosen tend to be the high-frequency ones the textbook taught first: "important", "because", "many people think". Those are the words a language model ranks as most probable, because they are the most common words in its training text.
Sentence shape is the second constraint. Learners are taught templates, and good learners use them: the topic sentence, the "firstly, secondly, finally" scaffold, the concluding "in conclusion". Native writers who have internalised the forms break them freely; learners still consolidating them keep to the pattern, because departing from it risks error. A learner also avoids the subordinate clause she is unsure of, so sentences come out shorter and more regular, which is the textbook description of low burstiness.
| Textual feature | Why a detector weights it | Why second-language writers show it |
|---|---|---|
| High-frequency vocabulary | Common words are the model's most probable predictions, lowering perplexity | Frequent words are acquired first and leaned on under time pressure |
| Formulaic connectives and frames | Stock phrases ("in conclusion", "on the other hand") are near-certain continuations | Explicitly taught in exam preparation and rewarded by rubrics |
| Even sentence length | Little variation in surprise across a passage reads as low burstiness | Long, risky constructions are avoided; fragments for effect are rare |
| Few idioms, little humour | Idiom and irony are high-surprise events that raise perplexity | Idiom is late-acquired and often avoided as unsafe in academic writing |
| Consistent register | Machines rarely shift tone, so a steady register matches the machine profile | One academic register is taught and held carefully |
| Correct but conservative grammar | Orthodox syntax is what a model predicts most confidently | Careful learners self-edit toward the safest structure they know |
None of these features is a fault in the writing. Several are exactly what a language course rewards, so the detector is penalising the student for having learned the lesson.
What Liang and colleagues found
The mechanism was demonstrated rather than merely argued in 2023. Liang and colleagues at Stanford ran 91 essays written by non-native speakers for the TOEFL examination, alongside essays by American eighth-grade students, through seven widely used detection services. On average the detectors labelled about 61 per cent of the TOEFL essays as machine generated, while the native-speaker essays were mostly passed as human; the study is published on arXiv as 2304.02819 and appeared in the journal Patterns the same year. The authors then confirmed the cause by intervening on it. When they asked a language model to rewrite the TOEFL essays with richer vocabulary, the false positive rate fell to roughly 12 per cent; when they simplified the eighth-graders' essays, those began to be flagged instead. Perplexity was driving the verdict, and it could be pushed either way.
The AI detector false positive problem is general
The non-native problem sits inside a general one. Weber-Wulff and colleagues, in the largest independent test of detection tools published in 2023, assessed fourteen services against human, machine, machine-translated and lightly paraphrased texts. None reached 80 per cent overall accuracy, and accuracy fell further once machine text had been paraphrased. Their conclusion was blunt: the tools are neither accurate nor reliable. This publication's earlier statement on assessment when the product no longer certifies the process reached the same place from the other direction: no inspection of a finished artifact restores the link to the mind that made it.
The equity consequence
International students pay the highest fees in most systems, are furthest from family support, and often hold visas that depend on continuous enrolment. Multilingual learners in schools are disproportionately from lower-income families. A tool whose errors concentrate on these students converts a technical weakness into an institutional bias, behind a number that looks objective. The teacher who acts on it intends no unfairness; the unfairness has been delegated to the software.
There is a chilling effect on learning itself. Students who hear that clean, careful prose gets flagged respond rationally: they roughen their writing, insert errors, or avoid writing in English where they can. The wider question of what integrity means once machines can write has to include the integrity of the institution's own procedures.
What a fair procedure looks like
Judgement is relocated rather than abandoned, onto evidence that actually speaks to authorship. The Initiative has argued since Collins, Brown and Holum's 1991 essay on making expert thinking visible that understanding shows in how work is built rather than in its finished surface. Applied to integrity, that principle yields three moves.
Ask for the trail. Where an assignment matters, require drafts, notes or version history alongside it, and treat their absence as a design failure of the assignment rather than a fault of the student. A tracked document records the sequence of composition in a form a second-language writer can produce as easily as anyone else.
Talk to the work. Five minutes on what the second paragraph is doing, why one source was preferred, and what was cut is more diagnostic than any probability score. Oral defence is also fair to learners of English, since it asks about ideas they hold rather than idioms they lack.
Compare with supervised samples. Keep a short piece of in-class writing from each student early in the term. When a later submission raises a question, the comparison is between that student's own supervised prose and the disputed text, one person rather than a statistical population. The non-native writer whose in-class work shows the same careful, formulaic character has been vindicated by her own hand.
Where a school continues to license a detector, the tool should generate a question, never a charge. A school writing this into policy will find a working template for an integrity policy in this lane, and a student on the receiving end of a flag will find what a student can show when falsely accused alongside it.
Frequently asked questions
Why do AI detectors flag ESL students more often?
Detectors score how predictable a text is to a language model. Writers in a second language draw on a smaller vocabulary, use the connectives they were taught, and keep sentence length regular to avoid error. Those habits lower perplexity and burstiness, the statistical profile machine text shows, so careful learner prose reads as machine prose.
What did the Stanford study on non-native writers find?
Liang and colleagues (2023) ran 91 TOEFL essays through seven detection services. On average about 61 per cent were labelled machine generated, while essays by American eighth-graders mostly passed. Rewriting the TOEFL essays with richer vocabulary cut the false positive rate to about 12 per cent, so word choice rather than authorship drove the result.
Is an AI detector score enough to accuse a student?
No. Independent testing by Weber-Wulff and colleagues (2023) found no tool above 80 per cent accuracy, with higher error rates for non-native writers and paraphrased text. A score is at most a prompt for a conversation. An accusation needs evidence about the student's process, and the student needs a fair chance to present it.
What should a teacher do when a detector flags an international student?
Treat the flag as a question. Ask the student to walk through the work, look at any draft or version history, and compare the essay with writing produced under supervision. If the process evidence is consistent with authorship, the matter ends there. If no process evidence exists, redesign the assignment rather than convict.
Where this leaves a school
The ai detector false positive is what the method produces when it meets a writer whose English is still being built, and no future version of the software changes the fact that predictability and authorship are different things. A school that understands this stops asking a probability score a question it cannot answer, and asks students instead to show how their work was made, the older test this Initiative has kept in view since its founding.
The practical path runs through the Academic Integrity hub, where the lane's articles on process evidence, disclosure and policy sit together, and through the guide to how accurate the detectors are in general, which places the non-native problem inside the fuller picture.