Skip to content
The 21st Century Learning Initiative

Writing to Think, When Machines Write First

School writing has never mainly been about producing text; it has been the discipline through which students learn to construct and refine thought. AI writing in education therefore poses a precise problem: when a machine supplies the first draft, it removes the struggle in which thinking was formed. Teachers must redesign assignments so the thinking remains the student's work.

What writing was doing all along

The research on composition settled a crucial point decades before machine text arrived: writing is not the transcription of finished thought but the activity in which thought gets finished. Flower and Hayes, in their 1981 cognitive process theory, described composing as a juggling of goals, plans, and revisions in which writers discover what they mean while trying to say it. Carl Bereiter and Marlene Scardamalia sharpened the developmental picture in The Psychology of Written Composition in 1987. Novice writers practise knowledge telling: they retrieve what they know about a topic and set it down in the order it arrives. Mature writers practise knowledge transforming: the demands of the text force them to reorganize, question, and extend what they know. The transformation is the education. A student who has wrestled a muddled idea into a defensible paragraph knows something at the end that they did not know at the beginning, and the knowing was produced by the wrestling.

This is why the Initiative's tradition treats writing as a laboratory case. The restored flagship on cognitive apprenticeship gives writing one of its three central examples: Scardamalia and Bereiter's procedural facilitation, which supplied students with prompts imitating the self-questions of expert writers. The intervention worked by making the invisible machinery of composition visible and practicable. Nothing about that machinery has changed. What has changed is that a student can now skip it.

The later instructional research points the same direction. Graham and Perin's 2007 meta-analysis of adolescent writing instruction, Writing Next, found the largest effects for explicit strategy instruction: teaching students the planning, drafting, and revising moves of skilled writers until those moves become their own. The finding matters here because strategy instruction is precisely what cannot be delegated. A machine can execute the strategies; only practice installs them in the student, and the installation was always the assignment's real purpose.

What the first draft was for

The first draft is where the cognitive work concentrates. It is slow, uncomfortable, and inefficient, and the learning science says those properties are not defects. Robert Bjork's account of desirable difficulties holds that conditions which impair immediate performance, generation, spacing, retrieval, often improve long-term learning; the generation effect in particular, documented since Slamecka and Graf's 1978 studies, shows that material a learner produces is remembered better than material merely read. Drafting is generation sustained for an hour. A meta-analysis of writing-to-learn interventions by Bangert-Drowns and colleagues in 2004 found modest but reliable gains in subject learning when students wrote about course content, an effect that depends entirely on the student doing the writing.

Machine text inverts the sequence. When an AI writing system produces the first draft, the student's task shifts from generating thought to appraising someone else's, a genuinely different cognitive activity and, for a novice, a much weaker one. Editing a competent draft never forces the confrontation with one's own confusion that drafting does, because the confusion has been papered over before the student could meet it. The essay arrives; the thinking never happened. This is the specific mechanism, worked out at the level of the whole model, in The Apprentice and the Answer Machine: assistance helps when it makes thinking more visible and harms when it makes thinking less necessary. First drafting is the least visible and most necessary thinking in a student's week.

The novice trap

There is an asymmetry in who can afford to delegate drafting, and it runs opposite to how the delegation actually distributes. An experienced writer who hands a routine passage to a machine retains everything needed to govern the result: a developed sense of the genre, a store of subject knowledge against which to test claims, an ear for what their own voice would and would not say. The machine saves them labor precisely because they no longer need the labor. A novice possesses none of that governing apparatus, because the apparatus is built by the very drafting being skipped. They cannot reliably see what is wrong, thin, or subtly off in the machine's draft; competent-looking prose forecloses the questions a struggling first attempt would have forced them to ask. So the student least equipped to supervise the machine is the one for whom supervision matters most, and the one most likely to accept the draft as given. Instructional research has a name for the general pattern: support calibrated for one level of expertise misfires at another. Here the misfire is systematic. The tool that legitimately serves the teacher's fluency undermines the student's acquisition of it, and a classroom policy that treats the two cases alike has misread the situation.

What machine drafting does not ruin

Honesty requires the other column. Not every act of school writing is writing to think. Some is transcription of decided content, some is formatting, some is the fifth rehearsal of a genre the student has already mastered. The research gives no reason to believe learning is lost when a machine handles writing the student was not learning from. There are also defensible uses inside the composing process itself: a machine can serve as a tireless generator of counterarguments, an instant supplier of contrasting introductions for comparison, a source of sentence-level alternatives a student must choose among and justify. Each of those uses puts the machine's text in front of the student's judgment rather than in place of it. The line is not between using and abstaining. It is between machine text as an object of the student's thinking and machine text as a replacement for it.

The same honesty applies to feedback, the other half of writing instruction. A machine's comments on a draft arrive instantly and in unlimited quantity, and some of them are good; the risk is not their quality but their timing and their authority. Feedback delivered before the student has finished struggling forecloses the struggle, and feedback from a source the student cannot argue with teaches deference rather than judgment. The research on formative response has always favored comments that pose questions over comments that supply fixes, and that distinction, not the source of the feedback, is the one worth policing. A teacher who routes machine feedback through the student's own evaluation, which of these five comments is right, which is wrong, and how do you know, has kept the judgment where it belongs.

Assignments that keep the thinking

The practical question is design. Assignments that survive machine drafting share a family resemblance: they make the process the deliverable. A teacher can require the draft trail, notes, outline, first draft, revisions, with the grade weighted toward the visible movement between versions rather than the polish of the last one. A teacher can set writing that depends on what only the student possesses: the particular discussion in Tuesday's class, the local data the student collected, the student's own earlier draft as the text under analysis. Short, frequent, low-stakes writing done in the room reclaims the generation effect directly. And where machine assistance is permitted, the assignment can require an account of it: what was asked, what came back, what the student kept, changed, or rejected, and why. That last move converts the machine from a shortcut into an occasion for exactly the metacognitive judgment Scardamalia and Bereiter were trying to provoke.

The pattern translates across subjects. In history, the assignment moves from "write an essay on the causes of the war" to "here are four sources we read this term; argue which one a machine's summary of the war most badly neglects, and show the summary." In science, the lab report's discussion section is drafted in the room, in fifteen minutes, from the group's own data table. In literature, the student annotates their own previous essay, identifying the two weakest claims and rebuilding them, a task for which no external draft exists to borrow. None of these designs is exotic, and each rests on the same principle: locate the writing where the thinking has to be.

What such assignments produce, incidentally, is evidence. When the process is the deliverable, the question of whether the student did the work mostly answers itself, which is why this essay's companion on assessment when product severs from process treats process design and integrity as one problem, not two. Teachers who want the full working context for these judgments will find it in AI Literacy for Educators: A Working Syllabus.

Implications for practice

Three commitments follow. First, decide, assignment by assignment, whether the writing is for thinking or for transcription, and protect the first kind: drafting that builds understanding should happen where the teacher can see it, in class, on paper, or in a tracked document. Second, grade the trajectory, not the artifact: ask students to show the distance between their first attempt and their last, and make that distance the thing that earns marks. Third, when machine text enters the room, keep it beneath the student's judgment: as material to critique, compare, and justify decisions about, never as an unexamined starting point. Writing to think survives machines that write first, but only in classrooms designed on purpose to preserve it.