Falsely Accused of Using AI: What a Student Can Show
The message usually arrives by email. An essay has been flagged, a meeting has been scheduled, and somewhere in the paragraph is a percentage: the detection service reports that the work is 87 percent likely to be machine generated. A student who is falsely accused of using AI reads that number as a verdict, and a parent reads it as a bill of charges. It is neither. It is the output of a statistical model that has never met the student, has not read the assignment brief, and cannot say which sentence it objects to or why.
This publication has argued for some time that a finished text is a poor witness to how it was made, and that the reliable evidence of authorship lives in the process: the drafts, the timestamps, the notes, and the student's ability to explain the work out loud. That position was developed for teachers. It applies with equal force to a student who must now defend writing they did in fact produce, and what follows is written for that student and their parent, with the teacher and the panel reading over their shoulder.
The short answer. A detector score is a probability estimate with published error rates above what any disciplinary standard should tolerate. Assemble your process evidence: document version history, earlier drafts and notes, browser and library history, and a written timeline. Ask the hearing to weigh it, request the detector report, and be ready to explain and extend your argument aloud.
What the accusation usually rests on
In most cases the entire case is the score. A teacher pasted the essay into a detection service, the service returned a percentage and a colour, and the percentage became the allegation. Sometimes there is a second strand: the writing seemed better, or flatter, or more formal than usual. That impression is a fair reason to ask questions and a poor reason to conclude anything, since students improve, imitate the sources they have just read, and write differently under different briefs.
A detector measures how predictable each word is given the words before it, and reports how closely that profile resembles text from a language model. It does not know who typed the words, cannot distinguish a student who writes plainly from a machine that does the same, and has no access to the only facts that matter in a misconduct case: what this student did, at what time, on which device.
Why a detector score is weak evidence
The research is settled enough to cite in a hearing. Liang and colleagues (2023) ran essays by non-native English speakers through seven widely used detectors and found that more than half were flagged as machine generated, while essays by native speakers were almost all passed. Sadasivan and colleagues (2023) showed that light paraphrasing defeats current detectors and argued that as language models improve, the best possible detector converges on a coin toss. Weber-Wulff and colleagues (2023) tested fourteen tools and found that none reached 80 percent accuracy.
A false positive rate in the low single digits, applied to every essay a school marks in a year, wrongly flags hundreds of honest pieces of work, and those it flags are disproportionately students who write in a second language, write plainly, or were taught a formal register. Our companion pieces on how accurate AI detectors are and why detectors flag non-native writers set out the studies; a student can print either and bring it to the meeting. None of this means the teacher acted in bad faith; a score is presented to teachers as a measurement, and they are rarely told its error profile. The task at the hearing is to replace a weak witness with better ones.
What a student falsely accused of using AI can assemble in an afternoon
The principle is the one this publication set out in Proof of Work: authorship is demonstrated by command of the process, and the process leaves traces. Collect them before anything is overwritten or deleted in a panic.
Start with the document. Google Docs keeps a version history under File, then Version history, with a timestamp on each entry; Word keeps one for files stored in OneDrive, and most learning platforms keep a submission log. A document composed over several sittings shows small additions, deletions and corrections spread across days, while a pasted machine draft appears as a single large insertion. Take dated screenshots, or better, ask the teacher to open the history with you: nothing persuades a panel faster than watching an essay grow on screen. The lane article on version history as evidence of authorship explains how to read the different shapes.
Then gather everything around the document: outlines, handwritten notes, a photograph of a planning page, messages to a friend about the topic. Browser history and library records show that the sources cited were actually opened, and when. Annotated readings show the reading that preceded the writing. Where a school has adopted a process portfolio, much of this already exists in one place.
Finally, prepare to speak. The oldest test of authorship is the oral one: an author can explain a choice, defend a claim, say what was cut and why, and take the argument a step further when asked. A student who wrote the essay can do this without rehearsal; a student who pasted it cannot. The lane essay on the oral exam as proof of learning describes what a fair oral check looks like.
| Evidence type | What it shows | Where to find it |
|---|---|---|
| Version history with timestamps | Composition over time, in sessions, with small edits | Google Docs (File, then Version history); Word files in OneDrive; the platform's submission log |
| Earlier drafts, outlines and notes | The argument developing before the final wording existed | Old files, notebooks, photos of planning pages |
| Browser and library history | Sources were opened, on dates before submission | Browser history, library loan records, database logs |
| Reading annotations | Engagement with the material before the writing | Marginal notes, highlighted PDFs, a reading journal |
| Supervised writing samples | The student's unaided style and level, for comparison | In-class essays, timed tests, earlier marked work |
| Oral explanation and extension | Command of the argument beyond possession of the text | Prepared for the hearing; can be offered unprompted |
| Written timeline | A dated account that ties the other evidence together | Written by the student, one page |
The timeline is one page in date order, from the day the assignment was set to the day it was submitted, with the evidence attached in the same order. A panel faced with a percentage on one side and a dated, cross-referenced account on the other is being asked to compare a guess with a record.
Asking the hearing to consider process evidence
Students are rarely told they may shape the procedure. Before the meeting, write to whoever convened it and make three requests. First, ask for the complete detector report: the tool used, the full output rather than a headline number, and which passages it flagged. Second, ask that the hearing consider the process evidence you will bring, and list it. Third, ask which policy is being applied and what standard of proof it sets. Put the requests in writing so the record shows they were made. A short statement, sent with the evidence or read at the start of the meeting, moves the conversation from a number to a record:
I understand that my essay has been flagged by an automated detection tool and that this is the basis for the concern. I wrote the essay myself and I want to help you confirm that. I am attaching the version history of the document, which shows it being written between [date] and [date] across [number] sessions, together with my notes, an earlier outline, and a timeline of the work. I am happy to discuss the essay in person. I would also ask to see the full detector report, and to know which policy and standard of proof the school is applying. Published research reports substantial false positive rates for these tools, particularly for writers whose style is plain or formal, and I would ask that the score be weighed alongside the evidence of how the work was produced.
A parent's role is to make sure these requests are made and answered, to attend if permitted, and to keep the meeting on evidence rather than character. Vouching for a child is natural and carries little weight; a timestamped draft carries a great deal.
What a fair procedure looks like
The burden of proof belongs to the institution. That is the standard in the misconduct codes of most universities, and there is no reason a school should hold a fifteen-year-old to a lower one. A fair procedure states the allegation and the evidence in advance, gives the student time to respond, allows a supporting adult to attend, considers evidence of process rather than product alone, applies a stated standard such as the balance of probabilities, records its reasons, and provides an appeal to someone who took no part in the original decision. A detector score alone should not meet that standard, and the lane's working template for an academic integrity policy for AI shows how to say so in writing.
What a fair panel wants is what a good teacher wants, which is to see the thinking. Collins, Brown and Holum's argument for making thinking visible was made about instruction in 1991, and it holds here: the way to know whether an apprentice can do the work is to watch the work, or, failing that, to examine its traces. There is a larger lesson in every false accusation, and it concerns homework and every other unsupervised product. If work was set where no process was visible, the school chose those conditions, and it cannot then treat the absence of process evidence as a mark against the student.
Frequently asked questions
Can a school punish me on an AI detector result alone?
It can attempt to, and some have, but it should not, and a written appeal that cites the published error rates and presents process evidence has a good prospect of success. Ask in writing which standard of proof applies, and insist that the score be treated as one item of evidence rather than as a finding.
How do I prove I didn't use AI if I have no version history?
Use what remains. Earlier drafts saved under other names, notes, outlines, messages about the essay, browser and library history and previous marked work all date the thinking. Offer to discuss the essay in person, and to write a short passage on the same topic under supervision so the panel can compare. Then turn on version history for every future piece of work.
Should I run my own essay through a detector to show it passes?
Usually no. Detectors disagree with one another, and with themselves on repeated runs, so a clean result from a second tool proves little more than the flagged result from the first. Time is better spent on the version history and the timeline. If you mention other detectors at all, use them only to show inconsistency.
Can my parents attend an academic misconduct hearing?
For students below university age, almost always. For university students, most codes allow a supporting person or a student union representative; ask when you request the hearing details. A supporting adult's most useful role is to take notes, make sure each written request is answered, and keep the discussion on evidence.
Where this leaves a school
A student falsely accused of using AI is carrying a burden the institution should have carried, and the way through it is to bring the evidence the detector could not see: the document's history, the drafts, the reading, and a live account of the argument. That evidence has established authorship for as long as apprentices have shown their work to masters, and it will remain good after the current generation of detectors has been replaced. For schools, the lesson runs the other way: any assessment that can be defended only by a score has been designed to produce accusations it cannot resolve. The Academic Integrity hub sets out the alternative, and version history as a routine part of every assignment is where a department should begin.