Complete model answers - 148 questions across all 21 chapters · Teaching English Writing to Generation Alpha
Generation Alpha refers to those born from approximately 2010 onward, with the cohort currently understood to extend to around 2025. McCrindle chose the term "Alpha", the first letter of the Greek alphabet, to signal a genuine new beginning rather than a continuation of the previous generational sequence. Having exhausted the Latin alphabet with Generation Z, the label marks this cohort as something qualitatively different: the first generation born entirely within the 21st century, and the first for whom digital technology was never an arrival but always simply the environment. The choice of "Alpha" is not a hyperbole. It is an acknowledgement that the frame of reference we used for all previous generations, adaptation to technolog, simply does not apply here.
For most teachers working with Gen Alpha in Gulf or urban Asian contexts, the most visible trait is the expectation of immediate feedback. Students submit a draft and check their phone within the hour. When feedback arrives four days later, they have mentally moved on. The writing task feels closed. The feedback feels irrelevant. The adjustment does not require responding to every draft instantly. It requires building one fast feedback loop into the drafting process itself. The simplest version: after students write a first paragraph, they swap with a partner for two minutes using one guided question: "Does the first sentence tell you exactly what this paragraph is about?" That immediate, low-stakes response keeps the student inside the writing problem long enough for revision to feel worthwhile. A second adjustment is AI-assisted self-assessment between drafts. Students paste their paragraph into an AI tool with a structured prompt and read the response before the next class. The teacher does not need to produce feedback on every draft, but the student receives a response while the writing is still warm.
The Parallel Attention paradox is the experience of a student appearing to be distracted such as when looking at a device, not making eye contact and who can nonetheless accurately repeat back what the teacher just said. It feels like multitasking but is more precisely described, following Kirschner and De Bruyckere (2017), as extremely rapid task-switching that feels simultaneous to the person doing it. The management implication is significant. If you build your device policy around the assumption that eye contact equals attention, you will make two kinds of errors: penalising students who are genuinely following while looking away, and missing students who are making eye contact while mentally elsewhere. The more reliable diagnostic is a brief comprehension check, a question, a quick written response, and a show-of-hands rather than visual monitoring. The practical shift is from "devices away" as a blanket policy to "comprehension check" as the actual measure. You may still decide that devices away is the right call for your context but the reason should be task focus and cognitive load, not the assumption that a student looking at a screen is not listening.
Because it changes what the teaching problem actually is. If a teacher believes their students do not write, the logical response is to build writing habits from scratch, to treat the class as starting from zero. That response misdiagnoses the situation and wastes enormous instructional time. The reality is that these students write constantly and fluently in Arabic, in Mandarin, in Spanish, in whatever language they live their lives in. They adjust tone, modulate register, manage audience, condense meaning, deploy humour. These are sophisticated communicative competencies. They exist. They are already practised. Recognising this shifts the teaching task from building a skill to transferring one. The teacher's job is not to instil the habit of writing or the concept of audience. It is to help students understand that the cognitive tools they already use in L1 apply in English too -- with different vocabulary, different grammatical conventions, and different genre norms. That reframe changes lesson design, changes the entry point for activities, and -- crucially -- changes how students see themselves. A student who understands they are already a writer, just in another language, approaches English writing differently from a student who believes they have never written anything of value.
They are two different sets of conventions for two different communicative contexts not two different levels of skill. The Digital Chat Toolkit is characterised by speed, informality, and shared context. Messages are short, often fragmented, sometimes abbreviated. Grammar and spelling are flexible. The reader is known and will fill in what is missing. Revision is rare because the medium is ephemeral. Feedback is instant. The Academic Writing Toolkit requires planning, structure, and explicitness. The reader is unknown or semi-known and cannot fill in gaps. Sentences must be complete and grammatically standard. The writer cannot rely on shared context. Revision is expected. Feedback is delayed. The critical point is that neither toolkit is superior. A student who writes with the Academic Writing Toolkit in a WhatsApp chat sounds strange and over-formal. A student who writes with the Digital Chat Toolkit in an essay sounds incomplete and careless. The skill is knowing which toolkit the context requires -- and that skill is already present in students' L1 repertoire. They switch registers effortlessly in Arabic or Tagalog. The task is to develop that same register-switching capacity in English.
Take one of the most commonly assigned tasks in EFL writing courses: "Write a paragraph about your hometown." From the student's perspective, the answer to "Why am I writing this?" is almost always: "Because the teacher asked me to." There is no reader beyond the teacher. There is no consequence if the paragraph is vague. There is no reason to be specific about the roundabout near the old souk or the smell of the Friday market, because no one reading the paragraph will care whether those details are accurate. One change transforms the task: add a reader and a reason. The revised brief might read: "A student from another country is coming to spend a month in your city. Write one paragraph that tells them one thing they would not find in a travel guide, something only someone who actually lives there would know." The topic is identical. The purpose is now real. The student suddenly has a reason to be specific, because vague information would fail the imagined reader. This adjustment takes thirty seconds to write into the task brief and produces measurably more invested writing. The audience does not need to be real. It needs to be believable.
The first sign is the blank page. A student who stares at an empty document after several minutes is not lazy or unmotivated -- they are overwhelmed. The gap between what they are being asked to produce and what they can currently manage is too wide to bridge alone. They have run out of scaffolding before they have started. The second sign is a large gap between oral fluency and written output. If a student can tell you, in conversation, a detailed and coherent story about their weekend but produces three halting sentences on the same topic when asked to write, the problem is not ideas or language. It is the cognitive load of the writing system itself: spelling, grammar, sentence construction, and meaning generation all at once. That student needs the cognitive demands separated. Let them speak first, record themselves, transcribe, and then edit. The oral fluency is the evidence that the ideas and the language are there. The writing instruction should help them access what they already have.
Writing is permanent in a way that speech is not. A spoken error disappears into the air. A written error sits on the page, visible, attributable, undeniable. For a learner who is already uncertain about their English, writing transforms every gap in knowledge into evidence. The blank page does not just ask "what do I want to say?" It asks "what am I capable of?" -- and for many learners, the honest answer feels inadequate. There is also an identity dimension that is easy to underestimate. In their first language, these students are full, capable, often eloquent communicators. In English, they are suddenly reduced to simple sentences, basic vocabulary, and constant error. That reduction is not just frustrating, it is disorienting. The person who appears on the page in English does not feel like them. Writing, unlike speaking, requires sitting with that inadequate version of yourself long enough to fill a page. This is why low-stakes writing (journals, ungraded drafts, choice of topic, simulated social media formats) is not a soft option. It is a deliberate reduction of the identity threat. The goal is to make enough space for the student to discover that their ideas are worth expressing before the grammar becomes the focus.
The ZPD is the gap between what a learner can do independently and what they can do with support. Vygotsky's (1978) insight was that this gap is not a deficit, it is the site where learning actually happens. Instruction that stays within what a student can already do produces no development. Instruction that reaches too far beyond current capacity also produces no development. The productive zone is in between: challenging enough to require support, close enough to current ability that the support can bridge the gap. Technology scaffolds the ZPD when it provides exactly the right amount of support, enough to help the student take the next step, not so much that it takes the step for them. A grammar checker that identifies an error and explains the rule is scaffolding the ZPD. A grammar checker that silently corrects the error is bypassing it. An AI tool that generates three possible topic sentences for a student to choose from and rewrite is scaffolding. An AI tool that writes the paragraph is substituting. The practical test for any technology tool is: does this help the student do something they could not quite do alone, in a way that builds their capacity to do it without the tool next time? If yes, it is scaffolding the ZPD. If the tool simply produces the output the student was supposed to produce, it is not scaffolding -- it is replacing.
This question is deliberately uncomfortable because the honest answer for many teachers is: "I do not actually know." Course design often proceeds from an assumption of adequate access, reliable home internet, personal smartphones, enough data to access cloud-based platforms outside school hours because it is easier to design for the average than to investigate the range. A simple first-week survey will reveal the actual landscape within ten minutes. In a class of thirty, there may be five students whose home access is genuinely limited, two or three who share a device with siblings, and several whose school-issued device cannot access certain platforms due to institutional filtering. The course design assumption most likely to need revisiting is the homework-as-digital-task default: assigning Google Doc submissions, Padlet brainstorming, or AI-assisted revision as out-of-class activities without confirming that every student can complete them outside school. The adjustment is not to abandon digital tasks, it is to ensure that every digitally assigned task has a paper backup built in, and that in-school device time is protected for students who need it. Access equity is not solved by assuming it exists. It is solved by designing around the students who have the least, not the most. PART I: FOUNDATIONS
Most practicing TEFL teachers trained during Wave Two or Wave Three, depending on their age and context. Teachers who trained in the 1990s were shaped primarily by the process approach arriving on the back of Wave One technology: multiple drafts, peer feedback, revision as a central skill. The permanent insight from that formation is that a single draft is not a finished piece of writing -- it is a starting point. That belief, once internalized, survives every subsequent technological change. It is still true in an AI classroom. Teachers who trained during Wave Three absorbed something different: the understanding that writing is social. Google Docs, wikis, and collaborative platforms made visible what was always true -- that writing develops through response, not in isolation. The permanent insight is that a student who never shows their writing to anyone except the teacher for a grade is missing the most powerful developmental mechanism available. Both insights are worth naming explicitly, because teachers who cannot articulate why they do what they do are more vulnerable to abandoning good practice when the next wave creates pressure to change everything.
The most universally reported stuck point across Gulf, Southeast Asian, and East Asian EFL contexts is the one-draft student: the learner who treats submission as the end of the writing process rather than the beginning of revision. The tells are consistent -- a draft submitted at the last possible moment, no evidence of revision between drafts, and a student who, when asked why they did not change anything after feedback, says some version of "I thought it was finished." The adjustment that works fastest is structural rather than instructional. Do not ask students to revise. Instead, make revision impossible to avoid by designing the assignment in stages: a planning submission due on Sunday, a first draft due on Wednesday, a revised draft due on Friday. Each stage is graded separately and carries weight. The Friday submission requires a short cover note answering one question: "What did you change between Wednesday and Friday, and why?" A student who submits identical drafts on Wednesday and Friday cannot answer that question. The structural requirement does what the instruction to "please revise" never does.
Take a common task: a descriptive paragraph about a place the student knows well, assessed on vocabulary range and sensory detail. On the Lean Tech Map, this task sits at the drafting stage. The stuck point for most students is not blank-page paralysis -- they know the place -- but lexical poverty: they have the ideas but not the English words for sensory experience. The tool most teachers reach for is Grammarly or a grammar checker, which addresses editing, not drafting. That is a stage mismatch. The simpler tool that actually addresses the stuck point is a vocabulary scaffold at the drafting stage: a sensory word bank generated in two minutes with a single AI prompt, or a pre-teaching activity using images of the place paired with vocabulary at the right CEFR level. Neither requires technology beyond a projector. The grammar checker can come later, at the editing stage, where it belongs. Identifying the stage mismatch -- using an editing tool at a drafting moment -- is the diagnostic the Lean Tech Map is designed to produce.
Most experienced TEFL teachers, if honest, sit closer to the accuracy end than the flow end -- not because they believe accuracy is more important, but because error correction is what they were trained to do, what observations and appraisals reward, and what feels most like visible, demonstrable teaching. Marking errors is active and documentable. Allowing imperfect drafts to develop is harder to explain to an observer who expects red ink. The question of whether this is right for students depends on stage. At A1 and A2, an obsessive focus on accuracy before fluency produces the worst possible outcome: students who write nothing, because nothing they can produce is correct enough to submit. At B2 and above, a complete focus on fluency without accuracy produces fossilized errors that become increasingly resistant to correction. The right balance shifts across a learner's development -- and the teacher whose accuracy-flow dial is stuck at the same setting for A1 and B2 students is almost certainly miscalibrated for at least one of them. The practical check: look at the last five writing assignments you graded. What proportion of your feedback addressed content and ideas versus grammar and surface errors? If the grammar feedback vastly outweighs the content feedback for A1 students, the dial may need adjusting.
AI is genuinely the simplest tool in three specific situations. First, when a student needs immediate feedback between drafts and the teacher is unavailable -- AI provides a response in seconds that a human cannot. Second, when a teacher needs to generate differentiated materials at multiple CEFR levels simultaneously -- AI produces a B1 version and a B2 version of the same task in the time it used to take to produce one. Third, when a student is stuck at the blank page on a topic where they lack content knowledge -- AI as a brainstorming partner surfaces possibilities the student would not have reached alone. AI is not the simplest tool when a peer conversation would do the job as well. It is not the simplest tool when a brief teacher conference would be more efficient. It is not the simplest tool for a student who lacks the language to evaluate AI feedback critically -- giving that student AI feedback without scaffolding the evaluation process is not scaffolding, it is offloading. The Lean Framework's insistence on simplicity is a check against reaching for AI by default. The question is always: is there a lower-tech solution that works just as well?
For most teachers, the honest answer is: partially. Knowing intellectually that previous generations of teachers panicked about word processors and then adapted does not fully neutralize the anxiety about AI, because the anxiety is not primarily intellectual. It is professional and relational. The fear is not really "will this technology destroy writing?" It is "will I still know what I am doing? Will my expertise still matter? Will my students still need me?" Those questions deserve honest answers rather than reassurance. The permanent insight from each wave is that the teacher's role does not disappear -- it shifts. Wave One shifted the role from transcription supervisor to revision coach. Wave Two shifted it from sole audience to authentic publishing guide. Wave Three shifted it from individual feedback provider to collaborative process facilitator. Wave Four is shifting it from writing instructor to thinking instructor -- the person who ensures that the ideas, the voice, and the critical judgment remain the student's own. That is not a smaller role. It may be a harder one.
The connection is direct. Each of the five stuck points is, in Vygotsky's terms, a description of where the ZPD is located for a particular student at a particular moment. The blank-page student's ZPD is at the pre-writing stage: they need support generating content, not producing polished prose. The one-draft student's ZPD is at revision: they need support seeing what revision is and why it matters. The repeated-error student's ZPD is at the editing stage: they need support understanding the rule, not just seeing the correction. The no-audience student's ZPD is at purpose and motivation: they need support imagining a reader before they can write for one. The perfectionist-shutdown student's ZPD is at permission: they need support understanding that imperfect drafts are not failures but starting points. The Lean Tech Map is, in effect, a ZPD map. It asks: at which stage is this student's development currently blocked, and what is the minimum support needed to get them to the next step? That is precisely Vygotsky's question, applied to a Gen Alpha writing classroom.
The framing matters enormously. A task designed to catch AI use -- where the explicit purpose is detection and prevention -- signals to students that the default assumption is dishonesty. That damages the relationship before a word is written. AI-resilient tasks are better framed around what makes writing genuinely personal: specificity, memory, local knowledge, and lived experience. A task that asks a student to describe their grandmother's kitchen in enough detail that a reader could draw it, or to explain a rule of a game they play with their family that has no English Wikipedia article, or to write about the moment they changed their mind about something -- none of these can be completed with AI because AI does not have access to those things. The student's specific memory and local knowledge are the content. The message to students is not "I do not trust you to write this yourself." It is "the most interesting writing comes from what only you know." That is true, it is pedagogically sound, and it happens to also produce tasks that AI cannot complete. The AI-resilience is a byproduct of good task design, not the purpose of it.
Independent production is where transfer happens. A student can analyze model texts, discuss revision strategies, and respond to AI feedback with apparent sophistication -- and still be unable to produce a coherent paragraph without support, because they have never been required to. The scaffold has become load-bearing. The ZPD analogy is useful here. A scaffold in construction is temporary. It holds the structure while it develops its own strength, then it comes down. A scaffold that stays up permanently is not a scaffold -- it is a crutch. Independent production in Stage 4 is the moment when the scaffolding comes down and the teacher discovers whether the structure can stand. In practice, Stage 4 looks like a short, in-class, timed writing task on a topic related to the lesson's focus -- no AI, no notes, no peer support. It does not need to be long. Three to five sentences are enough to reveal whether the structures practiced in Stages 1 to 3 have been internalized. The results are diagnostic as much as evaluative: a student who performs well in the scaffolded stages but struggles in Stage 4 is telling the teacher something important about where the ZPD actually is.
Authentic purpose produces better writing than artificial purpose. This principle predates every wave. It was true when students wrote on typewriters for teacher-only audiences and produced inert prose. It became demonstrable when Wave Two made real audiences possible and motivation visibly increased. It is still true in an AI classroom, where the students most likely to engage seriously with revision are the ones whose writing will be read by someone beyond the grade book. No technology changes this. Word processors did not change it. The internet did not change it. AI does not change it. A student who is writing to genuinely communicate something to someone they care about reaching will outwork, outrevise, and outperform a student writing to satisfy an assignment -- regardless of what tools either of them has access to. The implication for task design is the same in 1985 as it is now: before you decide what tool to use, decide who the writing is for and why it matters. Everything else follows from that. PART I: FOUNDATIONS
Table 3.1 sets out five stages: arrival, panic, early-adopter experimentation, consensus, and invisible infrastructure. The chapter states explicitly that AI is currently at Stage 2 moving into Stage 3 across most institutions - the panic stage, where assessment feels unsecured and professional expertise feels devalued, edging toward early adopters trying things without waiting for permission.
In a specific school, the evidence for Stage 2 usually looks like: no institutional policy yet, individual teachers making their own rules, and conversations dominated by concern about cheating rather than how to teach with the tool. Evidence of moving into Stage 3 looks like: a small number of teachers already experimenting openly, sharing what worked and what failed, without official guidance prompting them to do so. If a school can point to documented shared practice and a debate that has shifted from "should we" to "how do we", that is evidence of Stage 4 consensus forming, ahead of where most institutions currently sit.
Table 3.2 lays out three approaches that have dominated L2 writing instruction: product (1950s-1970s, writing as imitating correct models, marking errors in final drafts), process (1980s-1990s, writing as recursive drafting and revision), and genre (1990s-present, writing as varying by audience and purpose). The chapter is explicit that none of these is obsolete - effective teaching borrows from all three.
A teacher trained primarily in the product tradition will often still feel the pull toward a single final draft and toward marking errors as the main feedback mechanism, even while consciously building in revision stages. A teacher trained in process pedagogy will tend to build multiple drafts and conferencing into every task by default. A teacher trained in genre pedagogy will instinctively ask "who is this for and why" before assigning a topic. The diagnostic value of the question is in noticing the gap between which approach you were trained in and which approach your Gen Alpha students actually need - process and genre align with their digital habits far more naturally than product does.
The chapter's key insight from Wave One is that lowering the cost of revision increases the amount of revision students do. Before word processors, typewriters and handwriting punished revision with hours of labour, so students often submitted a first draft simply because producing a third draft was too much work. Once revision became a matter of keystrokes, multiple drafts became practical for the first time.
Most teachers apply this principle constantly without naming it: requiring digital drafts rather than handwritten ones, building in a revision stage between draft and final submission, or using "track changes" or comment tools to make revision visible. The chapter's caveat is worth sitting with - this wave is invisible infrastructure to Gen Alpha. They have never experienced the high labour cost of revision that made this insight necessary in the first place, which means the pedagogical value of teaching revision explicitly has not disappeared just because the technology that enabled it has become invisible to them.
Wave Two's permanent insight is that authentic audience produces better writing. Before the internet, a student's writing had essentially one reader: the teacher. The chapter cites Warschauer's (1996) finding that students writing for a real person who might actually respond invested in clarity and accuracy in ways that teacher-as-sole-reader tasks never produced.
Most regularly assigned writing tasks - essays submitted for a grade, paragraphs checked for homework - have the teacher as the only reader, even when the topic invites a wider audience. The simplest change suggested by the chapter's own examples is publication: a class blog, a shared Padlet visible to peers, or even just requiring students to read and respond to two classmates' work before submission. The point is not the specific platform - Gen Alpha finds email "formal and slow" - but the underlying principle that writing for someone who might genuinely respond changes how carefully a student writes.
The chapter's key lesson from Wave Three, learned by most teachers the hard way, is that putting students in a shared document does not automatically produce collaboration. Collaboration is a skill that requires explicit teaching - clear roles (writer, editor, reviewer, timekeeper), modelling of useful peer feedback rather than "good job" or "fix this," and active use of revision history to assess individual contribution.
A typical failure pattern: students are given a shared Google Doc and a group task with no further structure, one student does most of the writing, the others either don't engage or make superficial edits, and the teacher cannot tell who contributed what. Using Wave Three's insight, what was usually missing is the explicit structure - assigned roles, a model of what good peer feedback looks like, and a check of revision history before assuming collaboration happened just because the document was shared.
The chapter's TEFL Dimension caveat grounds this in Swain's Output Hypothesis (1985; 1995): producing language pushes learners to notice gaps in their competence, test grammatical hypotheses, and build fluency that input reception alone cannot develop. Writing in English, with all its difficulty and discomfort, is itself a language-learning activity for an L2 student - not just an academic exercise.
This is why the stakes are categorically different. In an L1 classroom, a student who uses AI to avoid writing loses a composition skill. In an L2 classroom, the same student loses a key mechanism supporting language acquisition itself. A student who outsources that difficulty to AI loses not just a grade opportunity but a development opportunity - which is why the chapter argues AI should increase, not decrease, the amount of meaningful English writing students produce, rather than simply being restricted or banned.
The chapter's Table 3.3 places Wave Four (AI) at "very high" relevance for Gen Alpha and describes their relationship to it as native: they expect machines to generate text on request, the way earlier generations expected a calculator to do arithmetic. Waves One and Two are invisible infrastructure to them, Wave Three is the default mode of writing, and Wave Four is simply expected.
Concretely, this means you cannot teach AI as a novel tool requiring justification, the way you might have introduced Google Docs to a class that had never used it. Your students already assume AI is available. The teaching task is not "should we use this" but "how do we use this without losing the language-production benefit that writing itself provides." Academic writing instruction has to build in deliberate friction at the right points: tasks designed so AI can support idea generation, vocabulary, or feedback, but cannot complete the actual thinking and language production the task is meant to develop. Because the expectation of AI access is already there, the burden is on task design, not on access control.
The chapter is explicit that this is the conviction the whole book is built around: AI should increase, not decrease, the amount of meaningful English writing students produce. The risk it identifies is substitution - AI writing the essay so the student produces none of the target language at all.
In practice, increasing meaningful output looks like using AI for the stages that are not the writing itself: generating topic ideas, producing model texts at the right level for analysis, giving instant grammar or organisation feedback so the student can revise more (and revise more often, the same Wave One insight about lowering the cost of revision), or acting as a low-stakes conversation partner for practice outside class. In each of these uses, the student is still the one producing the English sentences; AI is removing friction around that production, not replacing it. The test the chapter implies: after using AI, did the student write more English, or less? If a tool consistently results in less student-produced text, it is being used at the wrong stage of the task.
The chapter's closing key takeaway states it directly: "the teachers who navigated each wave well were not those who banned the tool or surrendered to it. They understood it well enough to use it purposefully." The attribute is not technical fluency with any particular tool - it is the willingness to understand a new technology deeply enough to make a deliberate pedagogical decision about it, rather than reacting with either prohibition or unconditional adoption.
Nada's own question to the professional development session - "is this different from all the other waves, or is it the same panic with better marketing?" - is itself an example of this attribute in action. She did not arrive with a verdict. She arrived with a structured question that let her place the new tool inside a pattern she already understood, which is what allowed her to respond to ChatGPT the same way she had responded to WordPerfect and Google Docs: with informed adaptation rather than panic or surrender.
Table 3.1 defines Stage 5 as the point where "the tool stops being technology and becomes simply how things are done," and where the panic of Stage 2 looks, in retrospect, disproportionate. The chapter notes AI is currently at Stage 2 moving into Stage 3, so Stage 5 is a forward projection, not yet a description of present classrooms.
Word processors are the chapter's own example of what Stage 5 looks like once a wave completes: nobody debates whether students should be "allowed" to use a word processor, the conversation moved decades ago from "should we" to "how do we teach revision well now that revision is cheap." Applied to AI, a Stage 5 classroom would not feature policies debating AI access; it would feature task design that simply assumes AI is present, the same way a lesson plan today assumes students can type. The teaching decisions that change are the ones currently consumed by the access debate - freeing that attention to focus on the actual pedagogical question the chapter raises throughout: which specific stuck points does AI address well, and which parts of the writing process still need to happen without it because the difficulty itself is the learning.
Most teachers who answer this question honestly find themselves at augmentation: they are using technology to improve a task that already existed, not to redesign it. Spell-checkers, digital handouts, and even Google Forms for quizzes are augmentation. They make existing tasks faster or cleaner, but they do not change what students are actually doing. The exercise is more productive when it is specific rather than general. Rather than asking 'where is my teaching on SAMR?', ask: 'where is this particular task?' Take one writing task you assigned last week. Was technology used at all? If not, that is substitution at best - the task assumed paper and nothing changed. If students typed instead of handwriting, that is still substitution. If spell-check was running, that is augmentation. If students left comments on each other's work in Google Docs, that is modification: the peer review process itself was redesigned. Moving one level higher is almost always a task-design question, not a tool question. The tool does not determine the SAMR level. The task does. Google Docs can be substitution (student types, teacher reads) or redefinition (student publishes to a global class blog, receives comments from readers in another country). The same tool, two different SAMR levels, entirely because of how the task was designed. The single most effective move from augmentation to modification is to make student thinking visible to other students. This usually means one of three things: shared drafts, visible peer comments, or a published artefact that someone beyond the teacher will actually read. Any of these moves the task beyond augmentation without requiring new tools.
This question is a diagnostic rather than a general reflection, and it works best when answered with a specific class in mind rather than 'students in general.' The toolkit addresses repeated errors most systematically, because LanguageTool provides immediate, rule-based feedback at the editing stage - exactly where error patterns need to be caught. The combination of LanguageTool flagging the pattern and a personal error log making it visible to the student over time is a complete response to this stuck point. Of the five, this is the one most directly matched to a specific tool with a clear mechanism of action. The stuck point most poorly addressed by the standard free toolkit is no sense of audience. Google Docs, Padlet, and LanguageTool are all tools that work within a closed classroom environment. They do not, by themselves, give students a reader who is not the teacher. Blogger and Google Sites can provide an authentic audience, but they require more setup and a deliberate decision to publish. Most teachers use the toolkit for the writing process stages - draft, revise, edit - without completing the cycle with genuine publication. The result is that the audience stuck point is addressed in theory (the toolkit contains publishing tools) but not in practice (those tools are rarely used). The practical implication: if no-sense-of-audience is visible in your current students, the fix is not a new tool. It is using a tool already in the toolkit - Blogger, Google Sites, or even a shared class Padlet - and treating publication as a required stage, not an optional extra.
The most common answer across EFL contexts is a dedicated peer review form - whether a Google Form, a printed checklist, or a separate feedback template. These are often used because teachers want structured peer feedback, which is a legitimate goal. But Google Docs' commenting and suggesting features already provide structured, location-specific feedback without requiring a separate tool. The comparison is instructive. A peer review form asks students to write general comments in a separate document: 'I liked how you...' and 'You could improve...' These comments are useful but disconnected from the text itself. Google Docs comments are attached to the exact sentence or word being discussed. The student can see the comment alongside the text, respond to it, and decide whether to accept the suggestion - all within the same document. Replacing the peer review form with Google Docs commenting removes one tool from the student's workflow, reduces the number of windows they need to have open, and actually improves the quality of feedback by anchoring it to specific text. The only loss is the structured prompts the form provided - which can be reproduced as a comment in the Google Doc itself ('leave at least one comment on the opening sentence, one on the argument, and one suggestion for the conclusion'). The broader principle: before adding any new tool, ask whether Google Docs can do the same job with a different setup. It can handle drafting, revision, peer feedback, teacher comments, version history, and publication to a limited audience. Starting with Google Docs and adding tools only when it cannot do what is needed keeps the toolkit lean and the cognitive load manageable.
The most common failure in technology-supported writing lessons is not tool malfunction - it is goal ambiguity. Teachers often choose a tool before they have clearly defined what students should be able to do by the end of the lesson. When the tool then fails or produces unexpected results, there is no clear benchmark for what to do instead, because the learning goal was never separated from the tool in the first place. The Saturday question that would have helped most in the majority of these cases is the first: what is the learning goal? Not 'what are students going to do with Padlet' but 'what will students be able to do with their writing that they could not do before?' If that question is answered clearly, the tool becomes interchangeable. Any tool - or no tool - that helps students reach that goal is acceptable. The backup question is the second most frequently underused. Teachers who have a paper backup ready before a lesson that depends on internet access lose nothing when the connection drops. They pivot to paper, complete the task, and reflect on whether the digital version would have added enough value to try again. Teachers without a backup lose the lesson and often the students' confidence in technology-supported work. A useful reframe: the four Saturday questions are not four equal checks. Goal and backup are the two that matter most. Tool choice is downstream of goal, and risk management is mostly solved by backup. If you only have time for two questions on Sunday evening, make them those two.
Across Gulf EFL secondary classrooms, the most commonly reported stuck point is the one-draft student - the learner who submits a first draft as a final product, does not revise between feedback cycles, and treats submission as the end of the writing process rather than the middle of it. Table 4.6 recommends Google Docs suggesting mode combined with an AI peer reviewer for this stuck point. Most teachers who report this as their primary problem are not using either of these tools. Instead, they are using one of two approaches: oral feedback in class ('I told them to revise and they did not') or written marginal comments on printed work ('I marked it and gave it back but the next draft looked the same'). The mismatch matters because both of these approaches leave the revision decision entirely to the student. There is no structural requirement to revise. Suggesting mode changes this: the suggestions are visible in the document itself, and the student must actively accept or reject each one. This creates a minimum revision interaction - the student cannot ignore the feedback without consciously dismissing it. The AI peer reviewer adds a second mechanism: three specific questions generated about the draft before submission give the student something to answer, which is a lower-demand entry into revision than the open-ended 'improve your paragraph.' The reason most teachers are not using these tools is usually setup time and familiarity, not disagreement with the approach. The practical response: teach suggesting mode in one lesson with no writing content - use it to edit a nonsense paragraph - so that students know how it works before it appears in a real writing task. Remove the setup barrier first, then use the tool.
Redefinition means the task enables something that was not possible before the technology existed. For a B2 student, this is achievable with currently available tools. A redefinition-level task at B2 might be: students write an opinion piece about a local issue, publish it to a class blog accessible to an international audience, respond to two real comments from readers outside the school, and revise their piece based on what those comments revealed about how their argument was understood. None of this was possible before the internet made real-time international publishing accessible. The technology is not decorating the task - it is constituting it. For an A1 student, redefinition requires more careful thinking. An A1 student cannot independently manage an international publishing workflow. But redefinition does not require complexity. A redefinition-level task at A1 might be: student records a voice note describing a picture using three sentences in English, sends it to a partner in a different class who draws what they heard, and the two then compare the picture and the description to find where language broke down. Voice recording, cross-class communication, and image-description comparison were not possible in a traditional classroom. The task is simple. The enabling technology is real. The learning objective - precise English description - is directly served. The principle that connects both: redefinition is about what the task enables, not how complex it is. A simple A1 task can be redefinition if it genuinely could not have been done before. A complex B2 task can be substitution if it is just a typed essay that used to be handwritten. SAMR level is determined by the task's relationship to the technology, not by the student's proficiency level or the sophistication of the output.
The CLEAR framework - Context, Length/Level/Limits, Example, Action, Role/Restriction - is introduced in [§ 5.3.1] with a full comparison table. The weak-vs-CLEAR prompt examples in that section demonstrate the principle this question applies: the difference between a weak and a CLEAR prompt is almost never about vocabulary or grammar. It is about specificity of purpose. A weak prompt - 'write a paragraph about my city' - tells the AI nothing about who is writing, at what level, for what purpose, with what constraints, or toward what standard. The AI fills all those unknowns with defaults: generic English, intermediate complexity, no particular structure, no particular voice. A CLEAR version of the same prompt specifies every one of those parameters. The output becomes directly usable as a teaching material. The ready-to-use teacher templates in [§ 5.3.2, Table 5.2] and the student-facing versions in [§ 5.3.3, Table 5.3] show what this specificity looks like at each writing stage. What this reveals: the quality gap between weak and CLEAR output is not a limitation of AI. It is a limitation of the instruction. Teaching students to write CLEAR prompts is therefore not only an AI literacy skill - it is a writing skill. The same precision that makes a good prompt makes a good topic sentence. The Student CLEAR Pocket Reminder in [§ 5.3.3] gives students a portable version of this principle.
The five-stage model - brainstorming, drafting, revising, editing, publishing - is laid out in [§ 5.4] with a separate subsection, sample prompt, and Teacher Caution box for each stage. The revision stage carries the highest risk, and [§ 5.4.3] identifies it explicitly: 'the most dangerous use of AI is paste paragraph, ask AI to improve, submit AI output.' The risk is higher at revision than at any other stage because revision is where the most meaningful learning happens. First drafts capture what students can already do. Revision is where they confront the gap between what they wrote and what they meant - precisely the gap that language acquisition lives in. When AI closes that gap automatically, the acquisition opportunity disappears. This connects directly to the flow-vs-accuracy tension discussed in Chapter 2 (§ 3.6 of that chapter) and to the over-reliance pitfall in [§ 5.6, Table 5.5] . The one structural change that most reliably reduces this risk is requiring a written explanation for every change made between drafts: for every sentence you changed, write one sentence explaining what you changed and why. This is the mechanism behind the AI Collaboration Log in [§ 5.8] which makes revision decisions visible to both student and teacher. The log cannot be completed by AI without revealing that the student did not understand their own revision.
The R element is introduced in [§ 5.3.1] as 'Who is the AI? What should it avoid?' and is explained in the Note on the L Element box in that section. The restriction component matters differently for L2 writers than for native speakers, and the Alpha L2 Intersection box in [§ 5.3.1] addresses this distinction directly. For native speaker contexts, restriction is largely a quality control issue. For L2 writing instruction, it is a learning protection mechanism. An AI without a restriction will do whatever produces the most coherent, complete, and polished response. If a student asks for help with a stuck sentence, AI will write a better sentence. If a student asks for revision feedback, AI will rewrite the paragraph. In each case, the student receives a completed product they did not produce. The restriction component - 'do not write the sentence, do not rewrite the paragraph, give options not answers' - is what converts an AI writing service into an AI writing scaffold. Without it, AI use in L2 writing instruction is almost inevitably acquisitionally counterproductive. The Teacher Caution boxes throughout [§ 5.4] all reinforce this: at brainstorming (§ 5.4.1), drafting (§ 5.4.2), revising (§ 5.4.3), and editing (§ 5.4.4), the caution is always the same - require the student to do the work, use the restriction to prevent AI from doing it for them. The sample prompts in each subsection model exactly how to phrase this restriction.
Activity 2 - AI as First Draft Buddy - is described in [§ 5.5] and is specifically designed to address the dependency concern through the memory-rewrite step. The criticism conflates two different kinds of dependency: cognitive dependency (the student cannot perform the task without AI) and process dependency (the student used AI as one input among several in a learning sequence). The second is not dependency any more than reading a model paragraph before writing is dependency. The memory-rewrite step is the critical pedagogical move. Closing the AI forces the student to hold the structure and some of the vocabulary in working memory and reconstruct it in their own language. This reconstruction is one of the central mechanisms of language acquisition: comprehensible input becomes output through the act of reformulation. The AI paragraph functions as comprehensible input of a kind the teacher could not provide at scale - personalised to the student's topic, targeted at their CEFR level. This connects to the ZPD scaffolding principle introduced in Chapter 1 (§ 1.9 of that chapter) and applied to AI use in [§ 5.5] . The reflection questions at the end of Activity 2 - what did you keep? what did you change? why? - are not optional extras. They are the mechanism that makes the activity acquisitionally valuable. A student who cannot answer them has not completed the activity. This is the same principle behind the AI Collaboration Log in [§ 5.8] applied at the activity level.
The AI Collaboration Log is introduced in [§ 5.8] and the CEFR-level scaffolding for it is explained in the Alpha L2 Intersection box in the same section. The student's objection is worth taking seriously rather than dismissing, because it identifies a real tension. The calculator analogy the student implicitly makes is addressed directly in [§ 5.7, Table 5.6] which distinguishes appropriate AI collaboration from inappropriate AI replacement. The response begins here: what is the learning objective of this writing task? If the objective is to produce a well-organised paragraph, then AI is no different from a calculator. But if the objective is to develop the ability to organise ideas in English - which is the actual objective - then using AI to produce the organisation is closer to having a calculator take the mathematics exam. The deeper issue the student's objection reveals is that they understand the writing task as a product task (produce a paragraph) rather than a process task (develop writing ability). This misreading is understandable in educational systems evaluated on products. The AI Collaboration Log is not surveillance of their honesty. It is evidence of their process, which is what is actually being assessed. The Class Agreement template in [§ 5.5] frames this for students directly: 'If you can answer these questions honestly, you are using AI as a tool. If you cannot, you are using it as a replacement.'
The distinction between using AI to generate teaching materials versus generating student work is the key principle here, and it is implied throughout [§ 5.3.2, Table 5.2] which provides CLEAR prompt templates designed for teacher use rather than student use. The most time-efficient application is generating differentiated practice materials. A single well-constructed CLEAR prompt can produce model paragraphs at three different CEFR levels for the same lesson in under two minutes - work that previously took forty-five minutes of teacher preparation. The prompt templates in [§ 5.3.2, Table 5.2] are ready to copy and customise for exactly this purpose. What makes this time-saving without reducing learning is the distinction between the teaching instrument (which AI generates) and the student output (which the student still produces). The cognitive work of the lesson still belongs to the student. The test for whether this saves time without reducing learning: would the teacher previously have done this preparation themselves? If yes, and if the AI output is reviewed and edited before use, the learning is unchanged and the time is recovered. The Lean Principle from Chapter 2 (§ 3.1 of that chapter) applies here too: one task, one tool, see what happens.
The vocabulary upgrade assignment is the clearest example in most EFL writing courses. It is consistent with the activities in [§ 5.4] and with the structured documentation approach in [§ 5.8] and directly addresses the Alpha L2 vocabulary gap described in the Alpha L2 Intersection boxes throughout the chapter. Students write a paragraph using their current vocabulary, then use AI to identify three instances of vague or overused words and request five alternatives for each. They choose one replacement per instance, write a sentence using it, and explain in one sentence why they chose it over the other four. This assignment is structured so that AI use is pedagogically unavoidable in the best sense: the student must first write independently, the AI then provides options - not answers - and the selection and explanation steps require active processing of exactly the kind that moves words from receptive to productive. At no point does AI write anything that appears in the student's final paragraph. The key structural feature that makes this work - the explanation requirement - is identical to the mechanism behind the AI Collaboration Log in [§ 5.8] applied at word level rather than draft level. A student who can explain why they chose one word over four others has actively processed all five words and understood the distinction.
The personal narrative is the clearest case for complete restriction, and the reason connects directly to the academic integrity framework in [§ 5.7, Table 5.6] and to the TEFL-specific concern raised in the Alpha L2 Intersection box in [§ 5.2] : 'A student who outsources that activity to AI is not just handing in work they did not do. They are removing the primary mechanism through which English develops.' A personal narrative asks a student to retrieve a specific memory, find the English words for specific sensory details and emotional states, sequence real events in a way that communicates their significance, and arrive at a reflection that is genuinely theirs. Every element depends on the student's actual experience. AI has none of those things. An AI-generated personal narrative is not a narrative in any meaningful sense - it is a simulation of the form without any of the content the form is supposed to carry. This makes it the clearest example of 'inappropriate' use in [§ 5.7, Table 5.6] : 'student did no writing or thinking.' The practical restriction uses the task-design principle from [§ 5.5] Activity 5 - AI Literacy: require two elements that AI cannot produce: a specific proper noun only the student would know, and an emoji reaction to the moment described, written before the draft, which the student then translates into words. Both are AI-proof because they require the student to reach into their own life and produce something irreducibly specific.
The distinction between appropriate collaboration and inappropriate replacement is formalised in [§ 5.7, Table 5.6] and is given a student-facing formulation in the Class Agreement template in [§ 5.5] : 'If you can answer these questions honestly, you are using AI as a tool. If you cannot, you are using it as a replacement.' This question asks you to construct your own explanation before using those resources. Begin with a question rather than a definition, because students remember what they figured out longer than what they were told. Ask: if you hire a translator to translate your essay into English, whose essay is it? Then ask: if you use a dictionary to look up one word and choose which fits your sentence, whose essay is it? Students will identify the distinction themselves: in the first, someone else made all the language decisions; in the second, they made all the language decisions and used a tool to help with one. This is exactly the logic behind [§ 5.7, Table 5.6] : the classification of a behaviour as appropriate or inappropriate turns on who is doing the thinking. The working rule that follows - can I explain why every sentence says what it says? - is the student-facing version of the AI Collaboration Log in [§ 5.8] and the Class Agreement in [§ 5.5] . For students who push back, the response is direct: you are not here to produce a good paragraph this week. You are here to become someone who can produce good paragraphs for the rest of your life. The Alpha L2 Intersection box in [§ 5.2] states this concern most forcefully: 'A student who outsources that activity to AI loses not just a grade opportunity but a development opportunity.'
The specific habits of Gen Alpha learners with vocabulary are described in [§ 6.7] under 'What Gen Alpha Already Does with Vocabulary', and the activation problem that results is introduced in the Alpha L2 Intersection box in [§ 6.1] . The Gen Alpha-Friendly Vocabulary Routine in [§ 6.7] is built directly on these existing habits rather than against them. The most consistent habit is asking a device. When students encounter an unknown word, they ask Siri, Google, or an AI chatbot rather than opening a dictionary. Building on this habit means redirecting the same instinct toward better tools. When a student is about to ask Google what a word means, the teaching intervention is not 'stop using your phone' - it is 'use this instead', with SKELL or Cambridge Dictionary (both listed in [§ 6.2.1] ) bookmarked on the same device. The habit of asking for immediate answers is preserved; what changes is where the answer comes from and what it contains. A second strong habit is code-switching. This can be leveraged explicitly: allow students to write a first draft in their L1 if they are stuck, then require them to translate and verify each sentence using Ludwig or SKELL before submitting. This positions L1 as a legitimate thinking tool rather than a failure state, while building the corpus-checking habit described in [§ 6.3.2] . The principle - speed without shortcuts - is stated directly in [§ 6.7] : 'give fast tools but always require a production step.'
The six pitfalls are catalogued in [§ 6.6, Table 6.5] with a corresponding solution for each. This question asks you to apply that table diagnostically to your own classroom. The most common pitfall across Gulf EFL contexts - and the hardest to address - is thesaurus overreach. Thesaurus overreach is harder to address than translation overuse because students feel they are doing something productive. Reaching for a more impressive word feels like effort. A student who writes exquisite instead of good believes they have improved their writing. The problem is invisible to them until the teacher marks it wrong. [§ 6.2.3] introduces the Thesaurus Rule - 'never use a word from a thesaurus that you cannot already define or use in a sentence of your own' - and recommends displaying it permanently in the classroom. That single rule gives students a self-check that costs no teacher time. The rule only works if it is explicitly taught, modelled, and returned to consistently. The classroom activities in [§ 6.2.3] - Which Word Is Wrong? and Thesaurus Challenge - are designed to make the rule visible and memorable before it needs to be enforced. The first time a student submits exquisite incorrectly, the feedback should reference the rule by name. After two or three referrals, students internalise the check without needing the teacher to enforce it.
The Gen Alpha-Friendly Vocabulary Routine is introduced in [§ 6.7] as a four-step daily structure: Discover (3 min), Define (3 min), Use (5 min), Review (4 min). The key design principle is that none of the four steps requires new materials if they are embedded into the opening of lessons already being planned. Discover uses whatever text is already on the agenda - students identify one unknown or uncertain word from the warm-up reading rather than from a separate source. Define uses a bookmarked dictionary or SKELL (listed in [§ 6.3.2, Table 6.2] ) - the teacher does not prepare the definition, the student finds it. Use produces one sentence the student writes in their notebook - the only teacher preparation is the instruction: 'write one sentence using this word that is true about your own life.' Review uses Quizlet or a simple repetition exercise - the Quizlet set can be student-maintained rather than teacher-prepared. The total teacher preparation time is zero if the Discover step is attached to existing lesson materials. The only investment is the first lesson in which the routine is taught, after which students run it themselves. The vocabulary-in-context principle behind the routine is consistent with the broader recommendation in [§ 6.5, Table 6.4] that vocabulary support should be integrated into every stage of the writing process rather than isolated as a separate activity.
The corpus verification step and the tools that enable it - SKELL, Ludwig.guru, and COCA - are described in [§ 6.3.2, Table 6.2] and demonstrated in the classroom activities in [§ 6.3.3] . The descriptive paragraph is the assignment where corpus verification produces the most immediate and visible improvement. Descriptive writing depends almost entirely on adjective-noun and verb-adverb collocations that L2 students get wrong in predictable ways. A student who writes strong rain, make a photo, or do a noise is not making random errors - they are applying L1 collocations to English. Activity 3 in [§ 6.3.3] - Fix the L1 Interference - is designed exactly for this: students search their L1-influenced combination in SKELL or Ludwig and find zero results, then search the correct English combination and find thousands. The corpus provides evidence, not just a rule, and as [§ 6.3.1] states: 'Students who see the evidence remember it far longer than students who are simply corrected.' The corpus verification step for a descriptive paragraph requires only one additional instruction at the revision stage - before you submit, choose two adjective-noun or verb-object combinations from your paragraph and check them in SKELL or Ludwig - adding approximately five minutes to the task. This maps onto the editing row of [§ 6.5, Table 6.4] where Ludwig.guru and SKELL are the recommended tools for checking if a sentence is native-like.
The receptive/productive distinction is introduced in [§ 6.1] and the specific shape this gap takes for Gen Alpha learners is explained in the Alpha L2 Intersection box in the same section: these students have substantial receptive vocabulary from screens but almost no practice producing English text. [§ 6.1, Table 6.1] maps the gap across vocabulary size, growth mechanisms, and technology strengths for each type. The gap appears most clearly in the disparity between what students understand and what they produce when writing about the same topic. A B1 student who reads and understands a paragraph about weather containing harsh, bitter, relentless, and overcast will write a paragraph about weather using cold, bad, and not nice. This is not a knowledge gap - the student knows harsh. They just do not reach for it when writing, because productive retrieval is a separate skill from receptive recognition and requires separate practice. Activity 4 - Word Replacement with Explanation - in [§ 6.4.3] targets this gap directly. The student is not taught a new word. They are required to retrieve words from their existing receptive vocabulary, select the most precise fit for the context, and explain that selection. The explanation step is the critical mechanism: it converts a passive encounter with a word into active processing of exactly the kind described in [§ 6.1] as what moves words from receptive to productive.
The framework the chapter provides for exactly this conversation is the receptive/productive distinction in [§ 6.1] and the Alpha L2 Intersection box in the same section. The student's claim is correct as far as it goes: they do know hundreds of words receptively. What the chapter's framework helps you show is that receptive knowledge is not the same as productive knowledge, and the gap between them is what vocabulary instruction is for. Begin by agreeing: yes, you do know hundreds of words. Then ask the student to write five sentences about a topic they know well, without looking anything up, and count how many different adjectives they used. Most will have used good, bad, nice, and big. Then ask them to read a paragraph on the same topic written at their level. They will recognise almost every word in it - but they did not write those words themselves. That gap - between the words you recognise and the words you reach for - is precisely what [§ 6.1, Table 6.1] maps as the difference between receptive size (larger) and productive size (smaller). YouTube and gaming grow receptive vocabulary quickly and are genuinely valuable for that - this is acknowledged in [§ 6.7] under 'What Gen Alpha Already Does with Vocabulary.' Productive vocabulary grows through a different mechanism: using words in sentences, checking collocations, making choices between alternatives, and writing the word enough times in enough contexts that it becomes automatic. The daily routine in [§ 6.7] is built around this mechanism: the Use step (writing one sentence) is what no amount of scrolling or gaming provides.
The two-source diagnostic framework is introduced in [§ 7.1, the Alpha L2 Intersection box] and provides the analytical tool this question applies. For teachers working with Arabic L1 writers, the most common sentence error is typically missing or incorrect articles - I met friend or The water is important for life of human. This is L1 interference, not a digital habit. Arabic does not use articles the way English does, and the gap is systematic rather than situational. The evidence that distinguishes source is consistency. [§ 7.1] makes this explicit: a digital habit error appears more in informal tasks and fades in careful academic writing. An L1 interference error appears consistently across all task types, including when the student is concentrating. Article errors persist regardless of task type. That consistency identifies the source as L1 interference. The diagnosis matters because the teaching response differs. [§ 7.1] states this directly: a digital habit error responds to register instruction, while an L1 interference error requires contrastive analysis and production practice. The common errors table in [§ 7.3, Table 7.2] maps error types to technology tools and approaches - using this table with your specific error in mind turns diagnosis into a concrete teaching plan. For Mandarin L1 writers, the equivalent diagnosis applies to tense marking: Mandarin does not mark verb tense morphologically, making Yesterday I go to the park an L1 interference error, not carelessness.
The diagnostic protocol is laid out step by step in [§ 7.6.2] and the problem with the Accept all approach is explained in [§ 7.6.1] . Most teachers, honestly assessed, are closer to Accept all than they would like - not through negligence but because running the full seven-step protocol for every sentence in a student paragraph is not feasible within a lesson. The most practical adjustment is selective application: choose the single most frequent error type across the class, tell students in advance that this lesson they will run the diagnostic protocol only for that error, and let everything else be corrected normally. [§ 7.6.3] models exactly this approach in the Grammar Detective activity, which targets one error type with a structured table. One error type, fully diagnosed, produces more durable learning than seven error types corrected automatically. A second practical move is to use the protocol as a whole-class activity before requiring individual application - the teacher projects one sentence, pastes it into LanguageTool together, and walks through the seven steps aloud. After two or three modelled sessions, students have the process in memory. The LanguageTool vs. Grammarly comparison in [§ 7.6.4] guides the choice of tool: LanguageTool is recommended as the starting point because it provides clearer explanations of why something is wrong, which is what makes the diagnostic protocol possible.
This experience is addressed directly in the Alpha L2 Intersection box in [§ 7.7] under 'Voice in a Second Language', which identifies exactly why this happens: the AI has more English than the student does, and an AI-expanded sentence often sounds closer to what the student means than what they can write independently. The temptation to copy is rational, not dishonest. The most recognisable version is the suddenly sophisticated paragraph: a student writing at A2 level submits a paragraph containing phrases like it is widely acknowledged that or this phenomenon can be attributed to a multitude of factors. The grammar is right. Nothing about it sounds like the student. What the sentence reveals is not dishonesty - it reveals that the gap between what the student wants to express and what their current English can carry is wide enough that AI output sounds more like their internal voice than their own writing does. The expansion protocol in [§ 7.7.1] addresses this by keeping the content entirely in the student's hands. The AI asks questions; the student answers in their own words about their own experience. The expanded sentence is built from what the student actually knows and feels, not from what AI generates about a generic topic. This is why the Alpha L2 Intersection box insists: 'the answer to the AI's question must come from your own life.' The Explain Your Change rule in [§ 7.8.3] adds a second layer of protection: students must document and explain every change, which is impossible if they have simply copied AI output they do not understand.
The four-level progression is introduced in [§ 7.4] with a dedicated subsection for each level (§ 7.4.1 through § 7.4.4), each with activities and technology support. Across most EFL contexts at B1 level, the hardest stage is not complex sentences but sentence variety. Students can produce a because clause when explicitly instructed to, but left unprompted they default to simple sentences for every paragraph. This is partly a digital writing habit - social media content is almost universally written in short, parallel structures - and partly a cognitive load issue: managing meaning, vocabulary, and grammar simultaneously leaves little capacity for structural experimentation. The sentence length audit from [§ 7.4.4] addresses this most efficiently: students paste a paragraph into Google Docs, identify all sentences under eight words, highlight them, and count them. Most students are surprised by what they see: a majority of their sentences are short and structurally identical. The visual makes the problem concrete. The task is then to combine three of the highlighted sentences with adjacent sentences using a conjunction or subordinating clause. This activity takes fifteen minutes, requires no new grammar instruction, and produces a paragraph that is visibly different from what students started with - the before-and-after visibility being particularly effective for Generation Alpha learners, as noted in the gamification section at [§ 7.11] . If the most difficult level for your students is compound rather than variety, the Conjunction Relay in [§ 7.10.3] provides the structured collaborative practice that addresses it.
The Three Attempts rule is introduced in [§ 7.8.2] as a prerequisite for AI use, and the step most frequently skipped is Attempt 2: use resources. Students move almost directly from writing independently to AI, bypassing LanguageTool, dictionaries, and corpus tools entirely. The reason is speed: these tools take longer than typing a question into ChatGPT, and Gen Alpha learners have been trained by their digital environment to prioritise the fastest available answer. The one change that strengthens Attempt 2 most reliably is making it visible and accountable. Add one column to the existing writing task template: Resource I used. Students must write the name of the tool or reference they consulted in Attempt 2 before they are allowed to reach for AI. This is consistent with the AI Collaboration Log introduced in Chapter 4 (§ 4.8), which documents the full AI interaction - the Resource column extends that documentation principle one step earlier, to the resource-use step that should precede AI. The column is graded for completion, not for correctness. Within two to three weeks, the habit of checking a resource before reaching for AI becomes automatic for most students. The secondary effect - noted in the technology tools table at [§ 7.9] - is that students begin to notice that LanguageTool and SKELL often answer the specific grammar and collocation question faster than AI does for targeted queries. For low-resource contexts, Attempt 2 can be operationalised as a printed grammar reference card or personal error log maintained by hand - the resource does not need to be digital, but it needs to be consulted and documented before AI is reached.
Sentence Surgery is introduced in [§ 7.11] as a classroom gamification activity that can run without technology. The adaptation to a fully no-device context is straightforward: the paragraph with errors is written on the board rather than projected from a screen, and teams record their diagnoses and fixes on paper rather than a shared document. Write a paragraph on the board with ten deliberate errors covering the types students are currently working on - the error categories in [§ 7.3, Table 7.2] and the pitfalls table in [§ 7.12, Table 7.5] both provide ready sources: two subject-verb agreement errors, two run-ons, two fragments, two word order problems, two dangling modifiers. For each error teams find, they must write three things: the sentence number, the name of the error type, and the corrected sentence. A correct diagnosis and fix earns one suture point. A diagnosis without a fix earns half a point. A fix without a diagnosis earns nothing - because naming the error is where the learning is. The paper version has one advantage the digital version does not: it forces team discussion before the answer is recorded. The conversation that happens between diagnosis and fix - where one student says I think it's a fragment but another says no, it's a dangling modifier - is precisely the kind of explicit grammar discussion that builds long-term awareness. That discussion happens less reliably when students work individually on a device. The competitive structure: the teams, the points, the reveal, and the prize - is the engagement mechanism, not the technology. Those elements transfer perfectly to paper, and in some classrooms, as [§ 7.11] implies in its no-tech gamification list, they work better there than on a screen. PART II: BUILDING BLOCKS OF WRITING
The trench coat problem is introduced in [§ 8.1] and its two possible sources - the digital writing habit and cognitive overload - are distinguished in the Alpha L2 Intersection box in [§ 8.1] . In most B1 Gulf and Southeast Asian EFL classrooms, the trench coat problem appears in roughly 40 to 60 percent of student paragraphs on the first task of a new unit, before any explicit paragraph instruction has been given. Two diagnostic tests distinguish the source. The first is L1 comparison: ask the student to write the same paragraph in their L1. If the L1 paragraph is also multi-topic, the source is the digital writing habit - the student does not organise by paragraph in any language. If the L1 paragraph is unified but the English one is not, the source is cognitive overload: the student knows how paragraphs work but cannot implement that knowledge while simultaneously managing English vocabulary and grammar. The second is task complexity: give the student a very simple English task where vocabulary and grammar demand is minimal. If the multi-topic problem appears on the simple task too, the source is the digital habit or the concept has not been taught. If it disappears on the simple task, the source is cognitive overload. The practical implication of the diagnosis is stated directly in [§ 8.1] : digital habit responds to explicit instruction about the concept of unity, backed by the colour-highlighting technique; cognitive overload responds to reduced task complexity and increased scaffolding (paragraph frames, L1 to English translation workflows) before raising the language demand. The paragraph frame progression in [§ 8.10] is designed precisely for the cognitive overload case: structure is externalised at Levels 1 and 2 so language is the only active variable.
Paragraph frames and templates are introduced across [§ 8.10.1 through § 8.10.4] in a four-level progression from full frames (A1 to A2) through templates with PEEL or PIE (A2 to B1) to independent writing (B1 to B2) to revision and feedback (B2+). The most common failure in classroom use of these frames is not introducing them too early - it is never removing them. The removal protocol that works most reliably is gradual and structured. In week one, students use a full frame with every sentence slot labelled. In week two, the labels are removed but the sentence slots remain. In week three, only the number of required sentences is specified. In week four, students write independently. At each stage, the teacher checks whether students can perform the previous stage without the frame before moving to the next one. A student who struggles at the no-frame stage goes back one step, not to the beginning. This is the scaffolding principle described in [§ 8.10] : the frame is temporary support, not permanent infrastructure. For teachers who have not used paragraph frames: the PIE template in [§ 8.10.2] and the PEEL template in the same section are the two models the chapter recommends. A B1 frame for a problem-solution paragraph would look like: My city has a problem with [topic]. The problem is serious because [reason 1] and [reason 2]. One solution would be [solution]. This would help because [benefit]. In conclusion, [summary of solution]. The which-model-to-use table in [§ 8.4.3] guides the choice between PIE and PEEL for your current students' level and task type.
The activity of feeding student paragraphs into AI for structural diagnosis is introduced in [§ 8.11.4 and § 8.12.2 through § 8.12.3] , and the important caveat about AI's generosity is stated in [§ 8.11.3] : 'AI-generated model paragraphs are often too polished,' and by extension, AI's diagnostic reading is often too charitable. Teachers who have tried this report a consistent finding: AI is more generous than they are. Where the teacher identifies a weak or missing topic sentence, AI often finds a plausible reading of an existing sentence as the topic sentence. Where the teacher identifies no supporting evidence, AI often reads an example as implicit evidence. This generosity is not a flaw in AI - it is a different kind of reading. AI is trying to interpret the paragraph charitably, looking for the best possible reading of ambiguous sentences. A teacher reading student work is asking whether the student has actually achieved the learning objective, which requires a stricter reading. The disagreement is productive precisely because it surfaces this distinction. The classroom activity in [§ 8.12.2] - Identify the Topic Sentence, Then Critique AI's Answer - is designed to generate exactly this disagreement and turn it into a teaching moment: when teacher and AI disagree about whether a topic sentence exists, the discussion should be: what would need to be true for AI to be right? Usually the answer is: the reader would have to do a lot of inferential work to find a topic sentence. Is that the reader's job or the writer's? This leads naturally to the key concept in [§ 8.3.1] : in academic writing, the topic sentence should be explicit, not implied.
The comparison between PEEL and PIE is in [§ 8.4.3, Table] which maps model choice to student level, task type, and difficulty. For most A2 and lower B1 students, PIE is the appropriate starting model. The three-part structure (point, illustration, explanation) is simple enough to hold in working memory while also managing English grammar and vocabulary. PEEL's four-part structure - with the link requiring both a summary and a forward transition - exceeds what most lower-intermediate students can manage on their first paragraph. The one lesson that moves the weakest students toward PIE most efficiently is the paragraph anatomy lesson: give students three correctly labelled PIE paragraphs in different genres (personal, academic, informal) and ask them to identify the Point, Illustration, and Explanation in each. This connects to the digital activities in [§ 8.5.1] - highlighting, matching, and critiquing - which provide the analysis-before-production sequence that reduces the cognitive demand of the initial writing attempt. Then give them three paragraphs where one element is missing. Students identify the missing element and write it in. Only then do students attempt their own PIE paragraph from scratch. This sequence shows students that PIE works across genres, preventing the misconception that it is a formula for one specific essay type. The paragraph frame in [§ 8.10.2] provides the structural scaffold for this first independent attempt, and the removal protocol described above (see Q2) applies from the first lesson onward.
Transitions are introduced in [§ 8.7.1, Table 8.4] , which catalogues common transition words by function. The knowledge vs. habit distinction this question applies is the same diagnostic logic used throughout the chapter for the trench coat problem and for paragraph frame removal - the teaching response depends on which problem is actually present. A quick diagnostic reveals which it is: give the student a gap-fill paragraph with blank spaces where transitions belong and a word bank containing the correct transitions. If they fill in the blanks correctly, the problem is habit: they know the words but default to not using them. If they fill in the blanks incorrectly or inconsistently, the problem is knowledge: they do not understand the functions the transitions serve. The Transition Word Relay activity in [§ 8.7] serves as both a diagnostic and a practice activity for the knowledge case. Habit problems respond to structural requirements: make submission impossible to complete without transitions. Require students to underline every transition word in their final draft before submitting. If there are no underlined words, the paragraph is returned. After two or three rounds of this, the habit begins to form. Knowledge problems require a function explanation, not a word bank - the student needs to understand that therefore signals a conclusion from evidence, that furthermore adds to a previous point. Without that understanding, transition words become decorative rather than functional. The tell is in the quality of transitions the student uses: randomly placed transitions with no relationship to the surrounding sentences indicate a knowledge problem. Consistently absent transitions in an otherwise coherent paragraph indicate a habit problem.
The colour-highlighting technique is introduced in the opening vignette and referenced throughout the chapter as the diagnostic method for the trench coat problem, most explicitly in [§ 8.13, Table 8.6] , which lists it as the technology tool for 'no paragraph unity.' The technique requires categorisation - assigning a marker to each topic - and any marking system that externalises that categorisation will produce the diagnostic benefit. The first no-device adaptation is pencil shading. Assign a different shading pattern to each topic: light horizontal lines for one topic, vertical lines for another, diagonal lines for a third, heavy shading for a fourth. Students mark up the paragraph with pencil patterns rather than colours. This requires no special materials beyond what is already on every desk. The cognitive work - categorising each sentence by topic - is identical to the digital version. The second adaptation is margin labelling. Students read the paragraph and write a one or two word label in the margin next to each sentence indicating its topic. Sentences that share a topic are then bracketed together. This takes slightly longer than colour-highlighting but produces more analytical engagement because the student must name the topic rather than simply recognise that a shift has occurred. The naming step often produces the insight independently: 'Oh, I wrote about three different things in that sentence - I can see it when I write it out.' The underlying principle - stated in [§ 8.13] and implicit throughout the chapter - is that the cognitive work is in the categorisation, not the colour. Any visible externalisation of that categorisation will produce the diagnostic benefit. PART III: ESSAY TYPES
The gap between oral storytelling ability and written narrative competence is the central problem this chapter addresses, stated most directly in the opening vignette (Taki) and in [§ 9.3] . The chapter identifies three consistent losses in the transfer from speech to writing: tense stability, descriptive specificity, and the point. Tense stability is lost because in oral storytelling, tense shifts are managed by performance - a speaker who switches to present tense creates immediacy, and the shift is interpretable from tone and gesture. In writing, the same shift looks like an error. The student who tells a compelling story in conversation and produces a tense-inconsistent written version has not lost their storytelling ability; they have lost the non-linguistic signals that make tense flexibility interpretable. This explains why past tense consistency is listed as a dedicated activity in [§ 9.8.5] : it is the most technically diagnosable of the three losses, and AI is particularly useful for it. Descriptive specificity drops because speaking involves implicit reference to shared context - 'we went to that place near the school' works in conversation because the listener knows what place is meant. In writing, the reader does not share that context. The student must specify, and specifying requires English vocabulary they may not yet have. Voice is lost because the student's storytelling voice lives in their L1, as the Alpha L2 Intersection box in [§ 9.3] states directly: 'His storytelling voice - the voice that hooks an audience and builds tension - lives in Arabic.' The transfer guide in [§ 9.4, Table 9.2] maps each oral/visual storytelling skill to its written equivalent and provides the bridge language students need.
This experience is addressed directly in [§ 9.3] and in the Alpha L2 Intersection box on 'The Voice Gap' in the same section. The generic narrative has always had the same tell: a narrative that describes events without any sense of a specific person experiencing them. The food was delicious. The scenery was beautiful. We were very happy. These sentences could describe any meal, any view, any happy occasion. They contain no detail that places a real person in a real moment. What makes a narrative feel authentic is specificity that cannot be faked. This is precisely why the chapter insists - in the second Alpha L2 Intersection box in [§ 9.10] - that tasks should 'require the story to contain details only the student can know: a specific smell from a specific kitchen, a specific phrase a specific person said, a specific feeling at a specific moment. AI cannot generate those things.' These details are irrevocable evidence of presence. A reader cannot receive them without feeling that the writer was actually there. The generic narrative existed before AI because students were performing the genre of personal narrative rather than actually remembering a personal experience - writing what a story about a family dinner is supposed to sound like, not what their family dinner actually was. AI has simply made this performance easier to sustain across a longer text. The practical response - offering genuine choice of story starter and requiring irreducibly personal content - is described in [§ 9.8.1] (Activity 1) and reinforced in [§ 9.10] : 'Personal story plus genuine choice is the combination that produces the most invested writing from Alpha L2 students.'
The principle of personal story plus genuine choice is introduced in the Alpha L2 Intersection box in [§ 9.10] and is the design logic behind the story starter activity in [§ 9.8.1] . Most standard narrative assignments offer choice of topic within a constrained genre: write about a memorable holiday, write about a challenge you overcame. The constraint means students who have not had memorable holidays or visible challenges are already excluded from the most natural interpretation of the task. The one change that most consistently introduces genuine choice is offering story starter options rather than topic categories. The difference is important. A topic (Write about a time you learned a lesson) presupposes the shape of the experience. A story starter (The time I changed my mind about something...) is open: students self-select the experience that fits, rather than trying to fit their experience to a predetermined shape. The CLEAR prompt in [§ 9.8.1] is designed to generate exactly these open, personal starters at the appropriate CEFR level. A second change is allowing students to propose their own story starter before the lesson begins, subject to teacher approval. This introduces genuine ownership before the first word is written. The chapter also notes that AI-generated starters can be generic - 'supplement with your own prompts based on what you know about your students' - which is a reminder that teacher knowledge of individual students' lives is itself a form of genuine choice provision that AI cannot replicate.
The six classroom activities span [§ 9.8.1 through § 9.8.6] , and Activity 3 - Narrative Structure with AI Scaffolding - is the best entry point for A1 to A2 students because it separates the thinking task from the writing task. The student fills in the narrative structure template in note form, not in sentences. What happened? Uncle house. We ate. There was football. This note-form planning is achievable at A1 level. The adaptation needed is to the template itself. The five-part structure in [§ 9.8.3] (beginning, problem, actions, resolution, lesson) is too cognitively demanding for A1. A three-part template works better: something happened, then something else happened, now I feel or I learned. Students write their three notes, then dictate the story aloud using voice typing, then read the transcript and choose five sentences to keep. This connects to the Lean Framework principle from Chapter 2 - the simplest tool that addresses the stuck point - applied to the narrative context. The AI component in Activity 3 (optional transition words) is removed entirely for A1. Instead, the teacher provides a printed transition word card (first, then, after that, finally) that students place on their desk and use as a reference while writing. This is consistent with the BYOD and limited resources guidance in Chapter 3 (§ 3.4): the resource does not need to be digital. For A2 students, the full five-part template can be used with note-form input, and the optional AI prompt for transitions becomes appropriate once students have completed the template.
The vocabulary expansion vs. content creation distinction is the central principle of [§ 9.7] , illustrated with the Good Example and Bad Example boxes in that section. The metaphor that works most directly for students is the cooking metaphor: you are the chef; AI is the supermarket. The supermarket gives you ingredients - words, phrases, options. You decide what to cook, how to combine the ingredients, and whether the dish is yours. If someone else cooks the meal, it is not your dish. You did not learn to cook. A more concrete test for students: could you explain every word in this sentence? If yes, it is yours. If there are words you cannot explain because AI produced them and you do not fully understand them, those sentences are not yours. This is the same principle as the Explain Your Change rule from Chapter 6 (§ 6.8.3), applied at the sentence level in narrative writing. The hardest case is when a student reads an AI sentence, understands it, and copies it because it is better than what they would write themselves - the rational response to the voice gap described in [§ 9.3] . The response: that is exactly why you are in this class. The gap between what you can say and what AI can say is what you are here to close. Copying AI does not close the gap. Writing your own imperfect sentences, getting feedback, and revising closes the gap. The imperfect sentence is the learning. This connects to the AI Collaboration Log from Chapter 4 (§ 4.8): if students document what they used and why, the distinction between tool and replacement becomes visible and self-regulating.
The so what? reflection is introduced as a narrative requirement in [§ 9.8.6] , which frames it as the lesson or takeaway that every narrative essay needs. Activity 6 in that section uses AI to help students identify what their story already implies - a useful scaffold, but one that still positions the reflection as something added at the end. The structural change that makes reflection unavoidable reverses this: make the reflection the first thing due, not the last. Before students begin writing the narrative, ask them to complete this sentence in writing: The thing I want the reader to understand from this story is... Students submit this sentence before the drafting begins. This does three things. It ensures the student has thought about the point of the story before writing it, which produces more focused narratives that naturally incorporate the elements in [§ 9.1, Table 9.1] - particularly the Point element. It gives the teacher a check against the final draft: does the story actually communicate what the student intended? And it makes the reflection an input to the writing rather than an afterthought appended to it. A student whose narrative does not match their pre-written intention statement has material for a genuine revision conversation: 'Your intention was X but your story seems to be saying Y. Which do you want to keep?' This conversation is more productive than 'you need to add a reflection sentence' because it treats the reflection as meaningful rather than structural. The Google Docs revision history spotlight in [§ 9.11] enables the teacher to verify this process across drafts: the revision history shows whether the reflection was present from the beginning or added at the last moment, which is itself diagnostic information about how the student understood the task. PART III: ESSAY TYPES
The priority order follows directly from the framework in [§ 10.1] and the Alpha L2 Intersection box in [§ 10.2] . In most A2 to B1 Gulf classrooms, vague vocabulary is the primary problem and it should be addressed first, because it is the root cause of at least two of the others. A student who writes the room was messy is not primarily failing to show rather than tell. They are failing because they do not have the English vocabulary to describe what mess looks like in specific terms. Clothes were piled on the chair comes from knowing the collocation piled on. If the student does not know those words, the showing technique taught in [§ 10.7] is inaccessible regardless of how well they understand the concept. The priority order: vocabulary first (addressed in [§ 10.5, Table 10.1] and the CLEAR prompt in [§ 10.4] ), then showing rather than telling ( [§ 10.7 and § 10.8] ), then spatial order ( [§ 10.11] ), then dominant impression ( [§ 10.13] ). Each requires the previous. A student cannot show without specific vocabulary. A student cannot create spatial order without showing-level specificity. A student cannot create a consistent dominant impression without spatially organised, specific details to accumulate.
The clearest evidence is the gap between oral description and written description - a phenomenon named directly in the Alpha L2 Intersection box in [§ 10.2] : these students are sophisticated visual consumers who have spent thousands of hours attending to colour, texture, atmosphere, and spatial arrangement in digital media. What they lack is not perceptual sophistication. What they lack is the English vocabulary to externalise it. Ask a student to describe a photo verbally and they produce rich, specific observations. Ask the same student to write a description of the same photo and they produce: the room is nice, the woman is there, it looks quiet. The oral description draws on perceptual sophistication the student genuinely has. The written description draws on English vocabulary the student genuinely does not yet have. The most practical response is the combination of voice typing and word banks described in the Tool Spotlight in [§ 10.15] : let the student speak the description in English and transcribe it, which often produces richer English than they would type directly, then upgrade the vocabulary using the sensory word bank generated through the CLEAR prompt in [§ 10.4] . The gap between what students can see and what they can write is not fixed. It closes with vocabulary instruction that is tied to their actual observation - which is exactly what [§ 10.6, Activity 1] is designed to provide.
The emoji-to-words transfer is introduced in the Alpha L2 Intersection box in [§ 10.16] and operationalised in the sample lesson warm-up in [§ 10.18, Table 10.7] . Teachers who have tried it report a consistent initial reaction: laughter, then genuine engagement. The laughter is the moment when students recognise that they are being asked to take their emoji use seriously as a communication system, which is both validating and surprising in a formal classroom context. The most productive emoji for the transfer exercise carry complex emotional and sensory information. Simple positive or negative emoji produce simpler descriptions. Complex emoji produce richer language. The technique works most powerfully for students whose oral English is noticeably richer than their written English - which is the Alpha L2 writer profile described in [§ 10.2] . Students who already have adequate written vocabulary may find the emoji step unnecessary. Students whose written vocabulary is far below their expressive range - the core population this chapter targets - respond most strongly. The Alpha L2 Intersection box in [§ 10.16] connects the technique to the vocabulary activation problem from Chapter 5: many of the sensory words students need are already in their receptive vocabulary from English-language media. The emoji prompt activates the connection between the known word and the observed experience. The technique does not teach new words. It surfaces words the student already partially knows and connects them to specific sensory experiences they have actually had.
The dominant impression concept is introduced in [§ 10.13] with a table of six examples (peaceful, chaotic, joyful, mysterious, lonely, nostalgic) each linked to specific detail types. The challenge at A2 level is that impression is an abstract noun that requires metalinguistic sophistication the student does not yet have. The one-feeling rule provides the accessible version: your whole description should make the reader feel one thing. If your reader feels peaceful, every detail should be peaceful. If your reader feels excited, every detail should be excited. If some details are peaceful and some are exciting, your description is confused. Pick one feeling before you start. Then only write details that belong to that feeling. The emoji-to-impression connection from [§ 10.16] makes this concrete: read your description and write an emoji next to each detail. If you have five different emoji, you have five different feelings. Remove or change the details that do not match the emoji you chose at the start. This gives the abstract concept an operational definition that works at A2 level: one emoji, all details match. Activity 5 in [§ 10.14] implements exactly this check: students read a partner's paragraph and identify the dominant impression it creates, then compare it to the writer's stated intention.
This question connects the vocabulary instruction in [§ 10.5 and § 10.6] to the assessment principles introduced in Chapter 12. The AI generates the word bank; the student selects from it. Three pieces of evidence are most reliable for confirming genuine selection. First, mismatched words. A student who genuinely chose from the word bank will occasionally choose a word that does not quite fit the sentence, because they chose based on partial understanding. A student who copied AI output will have consistently correct word usage because the AI generated the context around the word. Some imperfection in word choice is evidence of genuine selection rather than copying. Second, explanation. Ask the student: why did you choose this word rather than that one? A student who made a genuine choice can answer, even imperfectly. A student who copied AI output cannot answer because they did not make the choice. This connects directly to the oral defence activity in Chapter 12 (§ 12.13.4), applied here at word level: the explanation requirement converts a selection into an act of reasoning. Third, oral paraphrase. Read one of their descriptive sentences and ask them to describe the same thing without using the same words. A student who made genuine vocabulary choices can paraphrase because the words were connected to an observation they actually made. A student who copied AI output cannot paraphrase because the words were not connected to any personal observation. The AI Collaboration Log template in Chapter 12 (§ 12.6, Table 12.3) provides the written version of this same accountability mechanism.
The sample lesson in [§ 10.18, Table 10.7] has seven stages. The lesson survives without devices because the core instructional moves - emoji identification, word bank selection, sensory sentence writing, showing revision, dominant impression check - do not depend on technology. The warm-up (name three emoji for this place) can be done with any image printed or drawn on the board. Students write the emoji names in words rather than tapping actual emoji. The word bank generation, which the lesson assigns to AI, is replaced with a teacher-prepared printed card - exactly the kind of low-tech alternative implied by the BYOD guidance in Chapter 3 (§ 3.4). The teacher prepares sensory word bank cards for the most common descriptive topics (a room, a market, a meal, a street) and prints enough copies for the class. The cards function identically to AI output because the student's job is to select from available options, not to generate them. The only element genuinely hard to replace without technology is corpus verification of collocations - the SKELL check recommended in [§ 10.15, Table 10.5] . The low-tech alternative is a teacher-prepared collocation reference card listing 10 to 15 common collocations for adjectives students are most likely to use. The show rather than tell revision (§ 10.7) is done in pairs with pencil annotation: students circle every vague adjective and ask their partner: is this showing or telling? The dominant impression check is done orally. The underlying principle from the Lean Framework (Chapter 2) applies: the simplest tool that addresses the stuck point. For a no-device classroom, the simplest tool is a printed word bank card and a pencil. That is enough. PART III: ESSAY TYPES
The most reliable diagnostic is whether the student could answer a five-minute oral interview about their essay topic in their L1 without preparation - a check consistent with the oral defence protocol in Chapter 12 (§ 12.13.4). This principle is embedded in [§ 11.6, Activity 1] , which builds the knowledge check directly into the task: the teacher asks students to list things they know how to do or understand well before any topic is selected. Most standard expository tasks in EFL writing courses ask students to write about topics they do not genuinely know: climate change, nuclear reactors, historical events. These require research before writing, which immediately introduces AI as a shortcut. The Alpha L2 Intersection box in [§ 11.2] names the compound burden explicitly: students face two translation tasks simultaneously - visual to verbal AND L1 to English. Adding a third (unknown content to known content) makes AI dependency almost rational. The one change with the most reliable impact is constraining topics to the student's own direct experience. The CLEAR prompt in [§ 11.4] operationalises this by requiring topics that are relevant for a teenager and do not require expert knowledge. Activity 1 in [§ 11.6] provides the classroom procedure: students list things they know before any topic is selected, and those who are stuck use the CLEAR prompt to generate options they then filter by actual knowledge.
The visual-to-verbal translation problem is named in the Alpha L2 Intersection box in [§ 11.2] and illustrated in Faisal's vignette. The clearest evidence is the drawing test: ask students to draw the process they are about to write about, with arrows showing the sequence. Most produce a coherent, detailed drawing in five minutes. Then ask them to write the first paragraph. The drawing will almost always be richer - more steps, more detail, more accurate sequencing. The student knows more than they can write. The writing failure is not a knowledge failure. It is a translation failure. The activity that addresses this without technology is the annotated drawing sequence, consistent with the scaffold the intersection box recommends: let students first draw the process, annotate in their L1, then translate the annotations into English. Without devices, students draw the process on any piece of paper, label each step or arrow in their L1, then swap drawings with a partner. The partner reads the L1 annotations and writes an English equivalent below each one. The English labels become sentence starters for the process essay. This is consistent with the collaborative principle demonstrated in Activity 7 in [§ 11.13.2] : students check each other's structures rather than working alone. The annotation swap applies that principle to the pre-writing stage.
The most common pattern is fabricated statistics - a student writes a specific-sounding citation that does not exist. This is addressed directly in the Critical Caveat in [§ 11.1] and in Activity 3 in [§ 11.8] . The student is not usually trying to deceive. They are doing what they have been implicitly trained to do: add a specific statistic to sound credible. The caveat in [§ 11.7] reinforces this: if the final paragraph is identical to AI output with no changes, the student has not engaged with the task. The fact-checking activity in [§ 11.8] requires the student to verify at least two pieces of AI-provided information using reliable sources before writing their own evidence sentences. The section specifies what counts as reliable: a textbook, a trusted website such as NASA, BBC, or National Geographic. The student then cites the verified source, not the AI. The assessment mechanism that makes this durable is the Sources section requirement. Require students to submit a Sources section at the end of every expository essay listing the URL or book title for every fact they used. The teacher spot-checks two facts per essay. If either cannot be found in the cited source, the fact must be replaced or the essay is returned. The error table in [§ 11.12, Table 11.4] maps factual error directly to this protocol: require fact-checking with two reliable sources, cite sources. Consistency is the key: if fake sources are caught and addressed every time, students stop submitting them.
The 10-year-old test is introduced in Activity 4 in [§ 11.9] and returned to in Activity 8 in [§ 11.13.3] . The Alpha L2 Intersection box in [§ 11.13] connects this to the specific challenge for Alpha L2 writers: the default communication context for these students is shared-knowledge communication - texting friends who know the same things, posting for followers who share the same cultural references. Writing for a reader who knows nothing about the topic is a genuinely unfamiliar rhetorical situation. The most common pre-AI technique was the partner test: students read their explanation to a partner who genuinely knows nothing about the topic and count how many times the partner says what do you mean. This works but depends on the partner's willingness to be genuinely confused rather than politely nodding. Social dynamics in many EFL classrooms work against authentic confusion: students protect each other's face rather than providing honest feedback. AI changes what is possible primarily in terms of speed, patience, and availability. The 10-year-old simulation runs in seconds, can be repeated as many times as the student needs, and requires no preparation from the teacher. The intersection box states the principle directly: the student does not have to imagine the reader abstractly. The AI simulates the reader concretely, asking specific questions the student must answer. This reduces the cognitive demand of novice reader awareness to a manageable task - which is exactly what Alpha L2 writers need, given the three-way cognitive demand (content selection, language accuracy, vocabulary adequacy) they are already managing simultaneously.
The comparison and contrast essay is consistently the most difficult across B1 EFL contexts. The template in [§ 11.5.3] requires the writer to hold two subjects in mind simultaneously and organise the discussion around their relationship rather than around each subject separately. Students almost always default to the block structure (everything about A, then everything about B) because block organisation matches how they naturally think about two separate things. The specific sticking point is the conclusion. Students reach the while X and Y share some similarities, they differ in important ways sentence starter from [§ 11.5.3] and stop, because they have listed similarities and differences without deciding which difference matters most. The conclusion feels arbitrary because no evaluative thinking has preceded it. The fix is to add one pre-writing question to the template before students begin drafting: which difference is most important, and why? Students must answer this in one sentence before writing the conclusion. The AI support for this specific problem is mapped in [§ 11.12, Table 11.4] : ask AI to suggest a logical order for these steps and organise from most to least important. The Transition Word Wall tip in [§ 11.5] also addresses the comparison and contrast template specifically, providing the contrast and comparison signal words (similarly, likewise, however, on the other hand, in contrast) that students need to execute the point-by-point structure.
The core pedagogical move - separating the visual mental model from the language production task - requires no technology. This principle is stated in the Alpha L2 Intersection box in [§ 11.2] : the instructional response to the visual-to-verbal translation problem is to scaffold the visual-to-verbal step explicitly before asking for English. The drawing implements that scaffold without requiring any technology. The no-device sequence proceeds in four steps. Students draw the process on any piece of paper in five minutes - the drawing externalises their mental model of the sequence. Students then annotate each step in their L1, writing a word or short phrase. Students swap drawings with a partner. The partner reads the L1 annotations and writes an English equivalent below each one. The two students combined usually have more English than either has individually. Students then use their English annotations as sentence starters for the process essay, adding transition words from a classroom Transition Word Wall (tip in [§ 11.5] ). This sequence takes approximately 25 minutes and consistently produces more organised first drafts than students produce without the visual planning step. The partner annotation swap is the critical move: it transforms the translation step from an individual cognitive burden into a collaborative problem-solving task - consistent with the collaborative principle running through Activity 7 in [§ 11.13.2] and the broader toolkit approach in [§ 11.11] . Without devices, the drawing achieves the same structural visibility with paper and a pen. PART III: ESSAY TYPES
The diagnostic is direct and the chapter makes it explicit in the Alpha L2 Intersection box in [§ 12.2] : ask the student why they did not address the other side. If the student says I did not know how to write it, the problem is linguistic. If the student says there is no other side or the other side is just wrong, the problem is rhetorical. Both are common and require completely different responses. The linguistic problem responds to the sentence starters in [§ 12.5.2] . Give the student the counterargument and rebuttal template, show a completed example, and ask them to replicate the structure with their own content. Most students can do this within one lesson. The starters (Some people argue that... / However... / While it is true that..., this ignores...) provide the syntactic scaffolding the student is missing. The rhetorical problem requires a conversation that the template alone cannot produce. The most productive version is to show students two essays on the same topic - one that ignores the other side and one that acknowledges and rebuts it - and ask: which writer seems more knowledgeable? Which do you trust more? Students almost universally say the second. The Alpha L2 Intersection box in [§ 12.13] then provides the explanation: in English academic writing, presenting the opposing view fairly and then dismantling it is a sign of intellectual confidence. The counterargument and rebuttal sentence starters in [§ 12.5.2] address the linguistic layer. The intersection boxes address the rhetorical layer. Both are needed.
The pattern is predictable and consistent across Gulf, Southeast Asian, and East Asian EFL classrooms. The student writes: social media is bad because it makes people stupid. Also it wastes time. And it is dangerous. Therefore social media is bad. The position is stated three times with different words, which the student experiences as argument - exactly the behaviour described in [§ 12.2] under What They Already Do Well: state a strong opinion, repeat their position emphatically. The claim strength check from Activity 5 in [§ 12.10] would have been more efficient for one specific reason: it treats the weak claim as a design problem rather than an error. The student is not told their claim is wrong. They are told it is not strong enough to carry an argument, and given specific criteria: clear, debatable, and specific. The key insight stated in the section: the absence of evidence is almost always downstream of the vague claim. A student with a claim like social media is bad cannot provide evidence because the claim is too broad to evidence. A student with a claim like Instagram causes sleep disruption in teenagers because of its notification design can immediately ask: what evidence would prove this? The specificity of the claim generates the evidence question. [§ 12.12, Table 12.5] maps this directly: the fix for a weak claim is ask AI for three ways to make it stronger. The fix for no evidence flows from that, not the other way around.
The courtroom analogy is the most reliable approach with this age group and is consistent with the rhetorical explanation in the Alpha L2 Intersection box in [§ 12.13] . Imagine two lawyers in court. Lawyer A says: my client is innocent, do not believe the other side. Lawyer B says: the other side will tell you X, here is why X does not prove what they claim. Which lawyer do you trust more? Students almost universally say Lawyer B, because Lawyer B has clearly thought about the other side's argument. That is what counterargument does in an essay. The intersection box makes the cultural dimension explicit: in many Arabic, East Asian, and other rhetorical traditions, presenting the opposing view sympathetically before dismantling it is not standard practice. The student who says acknowledging the other side weakens my position is not making a logical error. They are applying a different rhetorical logic, one that is coherent within their own cultural and rhetorical context. The explanation that follows from this - different languages have different rules for how argument works, just as they have different grammar rules - is the one that lands without dismissing the student's cultural background. The sentence starters in [§ 12.5.2] provide the linguistic means. The cultural explanation in the intersection boxes provides the reason to use them. Both are necessary for the student who genuinely believes the convention weakens rather than strengthens their position.
The chapter's position is that it is a legitimate scaffold, and the reasoning is in the Alpha L2 Intersection box in [§ 12.13] . A student who reads three strong AI-generated counterarguments and then chooses which one to address has actually thought about the other side. That engagement is what the counterargument convention requires. The intersection box states this directly: AI is particularly useful here because it generates counterarguments the student would not have considered, forcing genuine engagement rather than a formulaic acknowledgement. Activity 3 in [§ 12.8] builds the distinction into the task design: students ask AI for the three strongest arguments against their position, choose one to address, and write a rebuttal using the sentence starters from [§ 12.5.2] . The AI generates the counterargument options. The student selects, evaluates, and rebuts. The thinking belongs to the student. The line runs through what the student does with the AI output. The condition that ensures genuine engagement rather than gaming is a selection explanation: students must write one sentence explaining why they chose this counterargument to address rather than one of the other two. This forces evaluation of argument strength. Activity 4 in [§ 12.9] extends this by requiring students to ask AI for feedback on whether their rebuttal logically responds to the counterargument - adding a verification step that keeps the reasoning in the student's hands. If the student can explain every significant choice, the scaffold is legitimate. If they cannot, the thinking has been bypassed.
The bandwagon fallacy (everyone knows that..., all students agree that..., most people think that...) is by far the most common across Gulf and Southeast Asian EFL classrooms. Its frequency reveals something specific about how these students conceptualise the purpose of argument - and connects directly to what [§ 12.2] identifies under What They Already Do Well: use emotional appeals, for example, everyone knows that. In informal arguing contexts, social consensus is a legitimate form of evidence. If everyone in a peer group agrees that a game is unfair, that consensus makes it functionally true. Bringing that logic into academic argument produces the bandwagon: the student cites the imagined consensus of an unnamed majority because, in their natural arguing environment, consensus is evidence. The Alpha L2 Intersection box in [§ 12.2] explains the compound source: monolingual Alpha learners transfer digital informal arguing habits into formal writing, while L2 writers may also be drawing on rhetorical traditions that weight collective social agreement differently from individual empirical evidence. Activity 8 in [§ 12.13.3] addresses this directly. Students paste a sentence containing everyone knows that into AI and ask: what logical fallacy is in this sentence, explain why it is a fallacy, suggest a better way to argue. [§ 12.12, Table 12.5] maps the bandwagon under emotional appeal without logic, with the fix: ask AI to rewrite this emotional argument as a logical argument, remove feelings, add reasons. The most efficient classroom demonstration: most people believed the earth was flat for thousands of years. Did that make it flat? Social consensus can be wrong. Academic argument requires evidence that is not dependent on what anyone believes.
The 90-minute lesson in [§ 12.14, Table 12.6] has eight stages. For 45 minutes, the split follows the logic of the lesson's design: keep the conceptually hardest work in class, move the production work to homework. Keep in class: warm-up (5 min), topic selection and pro/con list with AI using the CLEAR prompt in [§ 12.4] (10 min), claim writing with AI strength check from Activity 5 in [§ 12.10] (10 min), and counterargument and rebuttal practice from Activities 3 and 4 in [§ 12.8 and § 12.9] (15 min). These four stages cover the rhetorical and critical thinking dimensions that the Alpha L2 Intersection boxes in [§ 12.2 and § 12.13] identify as the compound challenge: choosing a position, making it specific enough to evidence, and genuinely engaging with the opposing view. They require teacher facilitation and cannot be approximated by homework. Move to homework: evidence gathering using the Activity 2 protocol in [§ 12.7] with the mandatory fact-checking step, full drafting using the five-paragraph template from [§ 12.5.1] , peer review using the argument checklist from Activity 6 in [§ 12.13.1] , and revision. The two-lesson structure that results - Lesson 1 covers thinking, between lessons students gather evidence and write a first draft, Lesson 2 covers revision - is actually more appropriate for Alpha learners than a single 90-minute block, consistent with the cognitive load principles discussed throughout Part II of the book. PART IV: ASSESSMENT AND PROFESSIONAL DEVELOPMENT
The key distinction: listing says what happened; explaining says why. The question that bridges them: 'why did that happen?' or 'what made that true?' If a student cannot answer that question verbally, they are not ready to write it. Oral rehearsal of causal reasoning before writing is one of the most effective scaffolds for this essay type.
Yes, it changes everything. A thinking error requires conceptual re-teaching. A transfer error requires pattern-specific correction and explicit contrast: in English, you choose one causal word, not two. The most effective response is to show the student the rule, give them two correct alternatives, and require them to choose. Not to mark it wrong and move on.
Most EFL curricula specify essay types but not topics. Within those constraints, genuine choice at the topic level is almost always possible. If the curriculum requires a specific topic, offer a menu of three: same cause-effect type, different subject matter. Any choice is better than none.
Two approaches: (1) pre-verify - teacher checks the AI's three most likely claims before the lesson and prepares one printed source for each; (2) delay verification - students flag uncertain claims with a question mark and verify as homework before submitting. Either approach embeds the critical literacy habit without requiring full internet access during class.
Activity 1 (The Arrow Game) is the best entry point: it requires no writing, it is fast, it makes the cause-effect distinction concrete and visual, and it generates classroom discussion about how students recognized the relationship. It also reveals immediately whether students have the basic conceptual understanding before writing is attempted.
Most students find starting from the effect more natural because the effect is often the visible, felt reality: exam failure, tiredness, stress. The cause is less immediately obvious, which is exactly what makes causal reasoning intellectually demanding. Starting from the effect and working backwards models how scientists, doctors, and analysts actually think. PART III: ESSAY TYPES
It is organizational. Tariq understands what comparison means. He does not know that a compare and contrast essay requires him to connect the two subjects within each paragraph. The question that reveals this: 'Tariq, in your second paragraph, can you show me the sentence where you connect online classes to in-person classes?' If he cannot point to it, he understands the problem immediately.
Drifting back to block structure is almost always a sign that the pre-writing stage was not explicit enough. Students revert to 'all about A, then all about B' because it feels more natural. The fix is to require students to write their three topic sentences before writing any body sentences. Topic sentences force the structure before the student can drift.
Sentence count is a proxy, not a perfect measure. It misses depth: a student might write three long, detailed sentences about A and four short, superficial sentences about B. However, sentence count is a reliable starting point for students who cannot yet perceive imbalance intuitively. The goal is to develop that intuition. Sentence count is the scaffold toward it.
Yes. Topics where the student has a very strong preference can produce advocacy rather than analysis. A student who strongly prefers TikTok over Instagram may find it difficult to give Instagram fair treatment. The correction: require students to write the body paragraphs first, without stating a preference, and add the preference statement only in the conclusion after they have given both subjects equal treatment.
Activity 1 (Venn Diagram) works entirely on paper. Activity 4 (Parallel Structure Clinic) works on paper with printed sentence pairs. Activity 5 (Padlet Gallery) can be replaced with a physical sticky-note board on the wall. Activities 2, 3, and 6 require the AI step to be moved to homework or a single shared device the teacher operates on the projector.
Most students do not know the transfer value of academic essay types. Telling them explicitly that a person who can write a well-organized comparison can write a product review, a job application comparing two offers, or a policy recommendation changes the perceived relevance of the task. Relevance is one of the four conditions most strongly associated with motivation in EFL writing research. PART III: ESSAY TYPES
Yes, with qualification. Academic writing distinguishes between personal anecdote as proof (not acceptable) and personal experience as illustration of a broader claim (acceptable and often necessary). Fatima's definition of freedom is valid in Layer 3 because it illustrates the concept, not because it proves it. The skill is learning when personal experience illuminates and when it distracts. That judgment is itself a B2 competency.
If definition writing is primarily cognitive, the most important intervention is in the pre-writing stage, not the drafting stage. Oral rehearsal of a definition before writing is more valuable than correcting a written definition after. Ask students to explain the concept aloud to a partner before they write. If they cannot explain it orally, they are not ready to write it.
It is invisible because students read for meaning, not for logical structure. They know what the sentence means, so it reads as correct. The most effective teaching moment is not correction but demonstration: show a deliberately circular definition of something unfamiliar. 'A glorbex is the thing that makes someone glorbexian.' Students immediately feel the emptiness. That feeling is what they need to recognize in their own writing.
Yes. AI-generated Layer 3 text is typically generic: 'in my personal experience, this concept has been important in my daily life.' The detection method: require at least one proper noun in Layer 3. A specific person's name, a specific place, a specific date or event. 'In my experience' is AI-compatible. 'In my mother's kitchen in Dammam on the day I chose what to cook for the first time' is not.
Yes. The three-layer framework works at B1 if Layer 2 is simplified: instead of dimensions and types, ask only for one thing the concept is not and one specific example. The significance statement can be a simple sentence: this matters because... What B1 students cannot yet do is the extended academic unpacking that Layer 2 requires at B2.
It is not a failure of writing. It is a structural failure. The most effective response acknowledges the quality directly: 'Yasser, your problem description is excellent. The detail about the exhaust smell in summer is exactly the kind of specific observation that makes a problem essay compelling. Now I need you to bring that same specificity to your solution. What, exactly, should be done, by whom, and when?' Acknowledgment first, redirect second.
The most common resistance is structural anxiety: students feel they cannot write solutions before they have fully described the problem. The response: the solution section does not depend on having read the problem section. It depends on knowing what the problem is, which students already know before they write anything. Reassure them that the problem section will be stronger for having been written after the solution, because they will know what it needs to set up.
The test adapts: Who becomes the person responsible for initiating change (which may be the student themselves). What becomes the specific behavior or action, not a feeling. When becomes a realistic timeframe for that behavior. 'I should be more patient' fails the test. 'I will pause for five seconds before responding in situations where I feel frustrated, starting this week' passes it. The principle is the same: specificity.
If AI's solutions are better, that is information. Ask students: why are AI's solutions more specific? What does AI know that you did not use? Usually the answer is that AI used generic knowledge efficiently. The follow-up task: take AI's most specific solution and localize it. Add one proper noun, one specific number, one concrete timeframe that only you can add from local knowledge. The localized version is always more persuasive.
The most effective reframe: 'Qualification is what distinguishes a thinker from a slogan-maker. Anyone can say this will definitely solve the problem. Only someone who has actually thought it through says this approach could realistically reduce the problem by half within two years. The second sentence is more powerful because it is more honest.' Show students a news article or policy document that uses qualified language. Ask: does the writer sound weak? Or credible?
Yes, and this is an important constraint. In some contexts, writing critically about government policy or institutional failures may feel risky. The solution is topic selection: offer a range of problem domains, from the very personal (time management, study habits) to the local and institutional (school canteen, transport) to the civic (environmental, infrastructural). Students choose the domain where they feel safe. The genre skills transfer across all domains. PART III: ESSAY TYPES
It comes from EFL instruction that prioritizes objectivity, accuracy, and deference to authority. It is particularly common in exam-oriented contexts where model answers are authoritative and deviation is penalized. The most effective response is structural: design a task that requires a position. When the assignment says 'write your opinion about...' and the rubric awards marks for a clear thesis, the permission is built into the task design.
The sequencing argument has merit, but it underestimates how much Gen Alpha students already know how to explain - they do it every day in voice notes and comment sections. The harder challenge is stating and committing to a position. Placing opinion writing early normalizes position-taking before the structural complexity of argumentative writing is introduced.
It is usually a thinking problem caused by a topic the student does not know enough about or care enough about. The fix is topic choice: if a student cannot generate three reasons, they have chosen the wrong topic. Allow them to switch. A student who genuinely holds a position about something they know well will generate reasons without difficulty.
It is a meaningful distinction at B1 level because managing a counterargument requires the student to hold two positions in mind simultaneously, support one, and rebut the other. That is a demanding cognitive and linguistic task. The distinction does not create a bad habit: it creates a foundation. Students who can write a consistent, supported position essay are much better prepared to introduce counterarguments at B2.
Activity 1 (the 30-Second Challenge) is the best entry point because it is oral, fast, and low-stakes. Students signal a position before they need to justify it. Activity 6 (the Padlet Opinion Wall) works for students who are reluctant face-to-face but comfortable expressing views digitally, which is true of most Gen Alpha learners. Anonymous posting removes the social risk of public position-taking.
AI has made it more common, not less. AI tools, when asked to write an opinion essay, frequently produce balanced non-essays because they are designed to present multiple perspectives and avoid taking controversial positions. Students who use AI to draft their essay often inherit this balance-without-commitment pattern. The AI-resistance strategy is thesis-first, personal experience-anchored writing: if the student writes the thesis and the personal example before touching AI, the position is established before AI can neutralize it. PART III: ESSAY TYPES
It is primarily a thinking error: Rania did not yet understand that a classification requires a single consistent principle. The most effective response is conceptual, not corrective: ask Rania the one question her teacher asked - are you classifying by what students do, or by something else? Once she identifies the principle, the categories follow naturally. Correcting the language without addressing the thinking produces better sentences with the same logical problem.
Classification is structurally demanding because all three constraints - single principle, mutual exclusivity, parallel structure - must hold simultaneously and visibly. Most other essay types allow partial compliance: a narrative essay with weak structure still reads as a narrative. A classification essay with overlapping categories or mixed principles fails as a classification regardless of how well it is written sentence by sentence.
Physical sorting works well with printed word cards or image cards. Students sort them on desks before writing. In a no-device context, the classification activity can be done entirely with pen and paper: students draw three columns, label them, and list members in each column. The Padlet gallery can be replaced with a physical board where students write their principle and categories on index cards.
The most effective transfer moment is before the task begins: 'This essay type is used by doctors to classify symptoms, by lawyers to classify types of cases, by biologists to classify species, and by teachers to classify types of learners. The principle you learn here - one criterion, no overlap, consistent structure - applies to every professional field.' Students who know the transfer value of a task engage with it differently.
The most widely agreed-upon decision is placing narrative and descriptive first: virtually all EFL teachers find that students engage most readily with personal and sensory writing. The most context-dependent decision is placing opinion before expository: in some curricula, expository writing is introduced much earlier. The decision that generates the most discussion is placing classification last: some teachers argue it should come after expository since it is a form of expository writing. Either placement is defensible; what matters is that the sequencing decision is explicit and principled. PART IV: ASSESSMENT AND PROFESSIONAL DEVELOPMENT
For most teachers who assign take-home essays with no process documentation requirement, the honest answer is: very little. The critical realization is that you cannot prevent AI access and you cannot reliably detect AI writing, so the assessment system must be redesigned around process evidence rather than product detection. The one change that creates the most visible trail is requiring Google Docs submission with revision history accessible to the teacher. A student who pasted AI output in one sitting has a revision history showing a single large text block appearing at once, followed by minor formatting changes. A student who wrote across three sessions shows incremental additions, deletions, and responses to comments spread across multiple timestamps. The second change that completes the trail is the AI Collaboration Log. A student who cannot produce a log documenting their AI interactions - what they asked, what the AI produced, what they changed, and why - has either not used AI (which the revision history supports) or has used it without documentation. Together, these two requirements make AI-generated submission visible without any detection technology, and without any accusation of dishonesty that cannot be evidenced.
The argument is well-grounded in second language acquisition theory. Writing in a second language is not only a composition exercise - it is a primary site of L2 output. When a student writes in English, they push their interlanguage to its limits, notice gaps between what they want to say and what they can say, and generate and test grammatical and lexical hypotheses. This is the mechanism Swain's Output Hypothesis identifies as a driver of acquisition. When AI writes for them, all of that cognitive work happens in a statistical model, not in the student's developing L2 system. The reframe is significant. If the primary concern is integrity, the response is disciplinary: this is cheating, here is the consequence. If the primary concern is acquisition, the response is different: you have lost the practice opportunity that would have helped you grow in English. The AI Collaboration Log supports this reframe structurally - it is not a policing mechanism but a learning artefact. A student who completes it honestly has documented their own thinking process, which is itself an acquisition activity. For Alpha L2 writers specifically, the reframe also acknowledges the genuine temptation: AI can produce more fluent English than these students currently can. The pull is rational, not dishonest. The response that acknowledges this - yes, AI's English is better than yours right now, and that is exactly why doing the writing yourself matters - is more likely to be heard than a punitive response that ignores the student's actual experience.
Teachers who have required AI Collaboration Logs consistently report three outcomes that are worth naming because they are counterintuitive. First, the frequency and sophistication of AI use decreases. Students who know they must explain what they did with AI output are less likely to paste it directly, because the documentation requirement makes the shortcut more effortful than it initially appears. The log functions as a friction mechanism without any policing. Second, students begin to distinguish between their own language and AI language in ways they had not previously done. The act of writing 'AI suggested this; I changed it because...' requires the student to read AI output critically and evaluate it against their own developing sense of their writing voice. This is exactly the metacognitive engagement the chapter argues for. Third, a minority of students initially document AI use dishonestly - logging that they did not use AI when they did. This is detectable by comparing the log to the essay and the Google Docs revision history. The act of requiring the log signals to the entire class that AI use is visible and documented, which changes the social conditions of the classroom even before any individual log is checked. For teachers introducing the log for the first time, the prediction should be: initial resistance, then gradual acceptance as students realise the log is lighter than they feared and the conversation it enables - about what AI is good for and what it cannot do - is genuinely interesting.
The distinction is precise and consequential. Corrective feedback tells the student what is wrong and provides the correct version: 'This should be However, not But.' Diagnostic feedback identifies the pattern behind the error and gives the student a tool to address it: 'You are using But where academic English expects a more formal adversative connector. Look at how Howeverand Nevertheless are used in the model text to see the difference in register.' Corrective feedback is faster to write and immediately satisfying - the essay looks better after the student applies it. But it transfers nothing. A student who changes But to However because the teacher marked it has learned that the teacher prefers However. They have not learned when adversative connectors are needed or how to choose between them. Diagnostic feedback takes longer to write but produces durable learning because it targets the underlying rule, not the surface error. The student who learns to ask 'is this register appropriate for academic writing?' has a transferable strategy that works on the next essay and the one after. The practical compromise most experienced teachers reach is selective diagnostic feedback: choose the one or two recurring error patterns that are most damaging to the student's writing - not the twelve things that are wrong - and diagnose those thoroughly. Leave surface corrections to LanguageTool and grammar checkers. Use your feedback time for the insights the technology cannot provide.
The practical obstacles are real and deserve honest acknowledgement. Time is the most significant: a genuine oral defence of even five minutes per student requires two and a half hours of teacher time for a class of thirty, before any preparation or recording. Scheduling is the second obstacle: finding time outside regular class hours in institutions with rigid timetabling is often impractical. Anxiety is the third: oral performance in a second language is significantly more stressful than written production for many students, particularly those from educational cultures that do not use oral assessment. The minimum version that preserves the core value of the oral defence is a three-minute spot-check on a random selection of five or six students per assignment cycle, conducted during class time. The teacher calls on a student without advance warning, asks two or three questions about their essay - 'Why did you choose this example? What does this sentence mean? What would you change if you wrote this again?' - and records a brief observation. A student who cannot answer basic questions about their own essay, in their own language if necessary, has not written it. This minimum version does not require additional scheduling, takes approximately twenty minutes per assignment cycle, and creates a classroom norm in which students know their essays may be discussed publicly. That norm alone is a significant AI-resistance mechanism, even for the students who are not selected.
The explanation that works most reliably uses an analogy the parent already understands: the difference between using a calculator in a mathematics lesson and having someone else complete the mathematics test. A student who uses a calculator to check arithmetic while solving a complex problem is using a tool to handle a low-level task so they can focus cognitive effort on the high-level task - the reasoning, the method, the interpretation. No one considers this cheating, because the mathematical thinking remains the student's own. A student who uses AI to generate their essay and submits it as their own has had someone else do the thinking. No English writing occurred. No vocabulary decisions were made. No sentence structures were developed. The student's L2 competence is exactly where it was before the submission - and in the L2 classroom, unlike the L1 classroom, that is not only an integrity issue but a learning loss. The policy in this course is that AI is a tool for specific, documented tasks at specific stages of the writing process - the same way a calculator is a tool for specific computational tasks. What the student may not do is submit AI-generated text as their own writing. The AI Collaboration Log makes the boundary visible: every use of AI is documented, and the final submitted text must be the student's own language. That is not cheating. It is learning with tools.
The honest answer for most EFL courses is: not enough, and sometimes none. Take-home essays dominate most EFL writing curricula, and timed in-class writing is often reserved for final examinations. This means that for many students, the first time they produce extended written English without AI access, without peer consultation, and without unlimited revision time is their high-stakes summative assessment. The case for more timed writing is not primarily about AI-resistance, though that is a genuine benefit. It is about developing the writing fluency that only practice under time pressure produces. A student who has written only take-home essays has no experience of the planning-under-pressure, first-draft-commitment, and revision-under-constraint skills that professional and academic writing regularly demands. Those skills require practice, not instruction. The minimum that serves both the acquisition goal and the AI-resistance goal is one short timed writing task per unit - fifteen to twenty minutes on a familiar topic with no preparation allowed - treated not as a test but as a formative writing sprint. Students submit immediately; teacher scans for patterns across the class; no individual grade is given. Over a course of twelve weeks, this builds a visible record of what the student can produce independently, which becomes the baseline against which take-home work is compared.
The response has two parts: acknowledging what is true in the student's claim, and explaining what the claim misses. What is true: professional writers do use AI, and increasingly so. Dismissing this as irrelevant would be dishonest and would lose the student's trust. What the claim misses: the analogy fails at a crucial point. Professional writers who use AI tools have already developed the underlying competence that the AI is assisting. A journalist who uses AI to draft a news summary has years of writing experience that allows them to evaluate, correct, revise, and take responsibility for what the AI produces. They are using AI to extend a capability they already have. A language learner who uses AI before developing that underlying competence is not extending a capability - they are bypassing the development process entirely. The AI Collaboration Log exists precisely to make this distinction concrete: it requires the student to evaluate AI output, decide what to use, what to change, and what to reject. That evaluation process is the competence being built. A student who accepts all AI output wholesale has skipped the evaluation step and produced nothing that develops their English. The Gradual Release Protocol does not remove AI support because AI support is wrong. It removes it because the goal of the course is not to make AI look good on the student's essays. It is to make the student capable of producing good writing - with or without AI - by the time they leave. PART IV: ASSESSMENT AND PROFESSIONAL DEVELOPMENT
The triggers for AI overwhelm among EFL teachers follow a consistent pattern: a sudden change in student behaviour (essays that are clearly AI-generated appear overnight), an institutional demand without guidance ('integrate AI into your lessons by next term'), a media event that makes AI sound either catastrophically dangerous or transformatively essential, or a tool update that renders a carefully planned lesson obsolete. What resolves it is almost always the same thing: a small, bounded experiment. A teacher who tries one AI activity with one class and observes what actually happens regains a sense of agency and proportion. The overwhelm is a response to perceived loss of control over the teaching environment; a single documented classroom experiment restores it by demonstrating that the teacher's professional judgment remains the most important variable in the room. For teachers for whom it has not resolved: the chapter's argument is that this is often because the response has been reactive rather than experimental. Consuming news about AI, attending panic-driven institutional briefings, and scanning social media for AI updates without trying anything in class increases anxiety without building competence. The antidote is action, not information. One CLEAR prompt, tried once, with one class, produces more useful professional knowledge than a hundred articles about what AI might eventually do to education.
Most EFL teachers' primary source of AI professional learning is reactive: social media posts, YouTube videos recommended by an algorithm, or a colleague who shares something interesting in a group chat. Reactive learning is fast and sometimes useful, but it has a structural problem: it is curated by engagement algorithms that favour novelty, controversy, and anxiety over accuracy and pedagogical relevance. The difference between reactive and systematic learning is what accumulates. Reactive learning produces a collection of unconnected observations with no framework to evaluate them against. Systematic learning - following two or three curated sources on a regular schedule - produces cumulative knowledge in which each new development can be placed in a context the teacher already understands. A teacher who follows one ELT-specific AI newsletter and reads one peer-reviewed article per month about technology in language teaching builds, over a year, a professional knowledge base that reactive consumption cannot replicate. The practical starting point is low-commitment: identify two sources worth following (one practitioner, one research-oriented), subscribe, and read one item from each per week. The investment is under thirty minutes a week. Over a term, the compounding effect of that consistent reading produces a teacher who can evaluate new AI tools against a principled framework rather than responding to each one as if it arrived from nowhere.
A CLEAR evaluation of a tool reveals, more often than not, that the problem is not with the tool but with the specificity of the instruction given to it. Teachers who evaluate a tool using CLEAR frequently discover that they have been using it at the wrong stage, for the wrong student level, or without adequate restriction - and that the mediocre results they have been getting are a direct consequence of vague prompting rather than tool limitations. The most commonly overlooked element is the R - Role and Restriction. A teacher using Gemini to generate vocabulary feedback for a B1 class who does not specify that the AI must use B1-appropriate metalanguage, must not exceed four observations, and must frame all feedback as a question rather than a correction, will receive output that is accurate but unusable with their class. The tool is not wrong; the instruction is incomplete. The evaluation also reveals stage mismatches. Many teachers who are dissatisfied with AI feedback tools are using them at the editing stage when the student's primary need is at the drafting stage. CLEAR's C (Context) forces the teacher to specify what stage the task is at, which often exposes the mismatch that was producing the unsatisfying results. The tool then used at the correct stage, with a specific Action and a clear Restriction, typically produces significantly better output - without any change in the tool itself.
The most useful single thing is not a tool recommendation, a framework, or an explanation of how AI works. It is a reframe of what the problem actually is. Most teachers who feel left behind believe the problem is that they do not know enough about AI. The chapter's argument is that this is the wrong diagnosis. The problem is not knowledge but agency: the feeling that AI is happening to them rather than being used by them. That feeling is resolved not by reading more about AI but by trying one thing with students and observing what happens. The five-minute conversation that works: 'Tell me one thing you currently do that takes a lot of your time and feels repetitive.' Whatever the answer - writing feedback on common errors, generating practice sentences, creating reading texts at the right level - there is a CLEAR prompt that addresses it. Write the prompt together in the conversation. Try it once. See what the output is. That first successful use, where the AI does something useful that saves real time, changes the teacher's relationship with the technology more than any amount of explanation. The chapter calls this the identify-try-reflect cycle. The colleague does not need to know the name. They just need to try one thing.
Solutionism in EFL AI integration most commonly appears as: using an AI grammar checker to address what is actually a motivation problem, deploying AI brainstorming tools for students who are stuck not because they lack ideas but because they lack topic knowledge, or introducing AI feedback tools for classes where the real issue is that students do not read feedback at all. The diagnostic question the chapter implies is: if AI were not available, what would I do to address this challenge? If the answer reveals a sound pedagogical approach - more peer discussion, more modelling, more time - then AI may accelerate that approach. If the answer is unclear, the problem is not yet defined well enough to choose a solution, and AI is likely to be applied to the symptom rather than the cause. Solutionism also appears in professional development contexts: teachers who adopt AI tools because the school requires it, before identifying what those tools are supposed to achieve. The chapter's Saturday night question - what does your student need to be able to do, and is AI the simplest thing that helps? - is the antidote to solutionism. It starts with the learning need and works forward to the tool, rather than starting with the tool and searching backwards for a use case.
Most teachers who ask their students about AI tool use for the first time report the same surprise: students are using a wider range of tools than the teacher assumed, including tools the teacher has never heard of. In Gulf contexts, students are frequently using AI tools embedded in Arabic-language platforms, or using AI through messaging apps, in ways that are entirely invisible to the teacher who is monitoring Google Docs activity. The second consistent finding is that students have developed their own heuristics for when to use AI and when not to - heuristics that are sometimes more sophisticated than teachers expect and sometimes more permissive. A student who says 'I use AI to check my grammar but not to write my ideas' has independently arrived at a version of the chapter's AI-resistance framework. A student who says 'I use AI for everything and then change a few words' has not. The conversation itself - structured as a discussion rather than an interrogation - produces three things: the teacher gets accurate intelligence about what is actually happening in the class; students feel their practices are taken seriously rather than policed; and the class develops a shared vocabulary for discussing AI use that makes the AI Collaboration Log more natural to complete honestly. The conversation costs fifteen minutes and produces information that shapes every subsequent lesson design decision.
The automation principle the chapter establishes is precise: automate the pattern identification, preserve the interpretation and response. This maps directly onto the distinction between corrective and diagnostic feedback. Correctional feedback on surface errors - article omissions, tense inconsistencies, comma splices, subject-verb agreement errors - is pattern identification work. LanguageTool and AI grammar agents do this reliably and quickly. The teacher's time is not needed for this layer. What AI cannot do is interpret why a particular error matters for this particular student at this particular stage of their development, connect the error to the student's L1 background, or decide which of fifteen flagged issues the student should focus on first. That sequencing and prioritisation is teacher judgment that produces the feedback the student can actually act on. The practical automation is therefore: run AI on the surface layer (grammar, spelling, punctuation) before the teacher reads the draft; receive a summary of the most frequent error types across the class; use that summary to write one class-level feedback comment that addresses the two or three most widespread issues; and reserve individual written feedback for the higher-order concerns - argumentation, organisation, voice, relevance to task - that the AI cannot evaluate. This typically reduces individual feedback time by forty to sixty percent without reducing the quality of the insight the feedback provides.
The productive middle is not a fixed point but a design decision that depends on the specific learning objective and the student's current stage. For fluency development at the drafting stage, immediate AI feedback on surface errors is appropriate. The student who is interrupted every few minutes by LanguageTool while writing loses flow; the student who submits a draft and receives a grammar summary an hour later has had their production protected and their accuracy addressed. The timing - during drafting versus after a first complete draft - makes a significant difference. For revision skill development, the right point is a delay of between twenty-four and seventy-two hours. A student who submits a first draft and receives AI feedback the same evening is less likely to have genuinely processed and internalised that feedback than a student who waits overnight before seeing it. The delay is not accidental inconvenience - it is a design choice that creates the conditions for genuine reflection. For the AI Collaboration Log specifically, the chapter's argument implies that immediate AI availability should always be paired with a mandatory reflection period before submission. The student who uses AI and immediately submits has not processed the AI's contribution. The student who uses AI, writes a Revision Plan, waits before drafting Draft 2, and then submits has engaged in exactly the kind of deliberate, reflective practice that produces durable learning.
The Gem that produces the most immediate value for most EFL writing teachers is the Error Diagnosis Gem, because it addresses the most time-consuming and most transfer-relevant form of feedback: identifying not just what is wrong in a student sentence but why, and providing a guiding question rather than a correction. The choice is strategic rather than urgent. The Vocabulary Preparation Gem saves time but primarily serves the teacher. The Lesson Plan Co-Designer Gem is useful but requires the teacher to already have a clear learning objective. The Rubric Score Reviewer Gem is valuable but requires calibration before it can be trusted. The Error Diagnosis Gem can be deployed immediately, requires no calibration, produces output the student receives directly, and builds the metalinguistic awareness that the chapter argues is the central long-term goal of sentence-level instruction. The restriction the chapter does not explicitly mention but that significantly improves the Gem's output is: specify the student's L1. An Error Diagnosis Gem instructed to 'remember that this student's L1 is Arabic and to flag likely Arabic interference errors specifically' produces diagnoses that name the L1 source of the error, which is the information that changes the student's response from 'I made a mistake' to 'my two languages work differently here.' That shift - from error-perception to language-contrast - is the difference between correction and acquisition. PART IV: ASSESSMENT AND PROFESSIONAL DEVELOPMENT
The broken system the chapter identifies has four defining features: feedback arrives late, covers everything, requires no action, and ends when the essay is submitted. Most AI feedback tools replicate three of the four: they cover everything (comprehensive automated reports on grammar, structure, vocabulary, and argument simultaneously), require no action (students read the report and submit the same draft), and end when the essay is submitted. The only feature they change is timing - feedback arrives faster, but the system's fundamental design is unchanged. A genuinely different system looks like the one described in the chapter: feedback is dimensionally focused (one dimension per agent, not everything at once), timed to enable revision (delivered before Draft 2, not after final submission), requires a structured response (the Revision Plan is not optional), and is connected to a second draft that is itself assessed. The multi-agent structure makes dimensional focus possible at scale. The Revision Plan requirement makes student engagement with feedback mandatory. The Draft 2 rubric score makes the connection between feedback and revision visible. What makes it genuinely different is not the AI. It is the cycle design. The same AI tools used in a submit-and-forget workflow produce the same broken outcome as any other feedback mechanism used that way. The chapter's argument is ultimately a claim about assessment design, not about AI capability: the system is the point, and the AI is the mechanism that makes the system feasible at classroom scale.
Both positions are partially right, and the disagreement is productive rather than resolvable by declaring one position correct. The colleague's position captures a genuine risk: feedback that is automated without a pedagogical rationale is a shortcut. A teacher who deploys the five-agent system because it saves time, without understanding how the feedback cycle is designed or how to monitor its outputs, has taken a shortcut. The monitoring protocol, the Revision Plan requirement, and the teacher's role in confirming recorded grades are not optional additions to the system - they are the pedagogical choices that distinguish automation from abdication. The chapter's position captures what the colleague's position misses: the feedback loop is not removed by automation. It is restructured. The teacher is removed from first-pass surface error marking, which is pattern identification work that AI does reliably. The teacher remains essential for monitoring trends, intervening with struggling students, and confirming grades - work that requires professional judgment that AI cannot provide. The evaluative question is: which tasks in the feedback loop require teacher judgment, and which are pattern identification that can be reliably automated? The chapter's answer is specific: grammar, vocabulary, macro-structure, and task achievement can be assessed dimensionally by constrained agents with acceptable reliability. Evaluating the appropriateness of a student's voice, the depth of their engagement with the task prompt, or the significance of their argument requires a reader who knows the student and the context. That reader is always the teacher.
The advantages are diagnostic and motivational. A holistic score of 12 out of 20 tells the student they are below proficient. A dimensional profile of 3-3-2-1-2 tells the student that their vocabulary and grammar are solid, their macro-structure is developing, and their micro-organisation and task achievement need the most attention. The second message generates a revision priority that the first message cannot. The chapter's claim that the system's most important output is better writers rather than better essays depends on this diagnostic specificity: a student can only improve systematically if they know which dimension to work on. The motivational advantage is related: a student who sees that three of their five dimensions are at 3 or above is more likely to engage with feedback on the weaker dimensions than a student who receives an undifferentiated low grade. Partial success is visible in the dimensional profile in a way that holistic scoring conceals. The risk the chapter does not discuss explicitly is the risk of gaming. Once students understand the rubric, they may write strategically for the dimensions rather than for genuine communicative competence - producing essays that are structurally correct and grammatically clean but intellectually thin. This is a known problem in standardised assessment research. The Revision Plan requirement partially addresses it, because a student who has gamed the rubric cannot write a genuine Revision Plan about how AI feedback changed their thinking. But the gaming risk is real and teachers using the system should watch for essays that score well on all five dimensions but feel empty.
The case has three parts, and the most persuasive for students is the third. First: the Revision Plan is the evidence that the feedback was read. A student who cannot write 100 words about what the feedback said and what they changed as a result did not engage with the feedback. The plan is not additional work - it is the minimum demonstration that the time the feedback took to generate was not wasted. Second: the plan forces the student to prioritise. The agents provide specific observations on five dimensions. The Revision Plan asks the student to decide what matters most and to commit to two or three specific changes. That prioritisation is itself a writing and thinking skill - the same skill that distinguishes a competent reviser from a student who makes random changes between drafts and calls it revision. Third, and most persuasive: the Revision Plan is the part of the assignment most resistant to AI interference. A student can generate a draft with AI. They cannot generate a genuine Revision Plan about their own feedback experience - not one that accurately describes what the Grammar Agent said about their specific sentences, what they decided to change, and what they chose not to change and why. That specificity requires genuine engagement with the feedback report. For students who want to demonstrate that the work is their own, the Revision Plan is the most reliable place to do it.
In Gulf EFL contexts with large classes and institutional pressure to produce grade results quickly, the highest risk failure condition is almost always the same: the system is deployed without teaching the rubric first. When students receive dimensional scores without understanding what each dimension means or how each score level is described, the feedback report is an oracle - authoritative but uninterpretable. A student who receives a 2 on Task Achievement and does not understand what Task Achievement means in the context of the rubric cannot act on that score. They experience it as judgment, not guidance. The revision that follows is often cosmetic rather than substantive, because the student is guessing at what the score means rather than knowing. The design choice that most reliably prevents this is built into the chapter's recommended sequence: before the first submission to the system, students score two sample essays using the rubric. Not in pairs, not as homework, but as a whole-class calibration exercise where the teacher projects the sample essay and the rubric, students score each dimension individually, then scores are compared and discussed. After that exercise, students know what a 3 on Grammar looks like, what distinguishes a 2 from a 3 on Task Achievement, and why the Synthesiser's score on their own essay is what it is. The fifteen minutes that exercise takes is the highest-return investment in the entire implementation.
The research ethics obligations are substantial and the chapter is explicit about them in Section 21.8. They deserve to be stated as conditions rather than suggestions. Before any student data generated by the system is used for any purpose beyond delivering feedback to that student, five conditions must be met. First: institutional approval from whatever body governs research in your context - school committee, university ethics board, or ministry-level approval if required. Second: informed parental consent for minor students, in the family's home language, making clear that participation is voluntary, that refusal carries no academic penalty, and that feedback delivery is identical whether or not the data is used for research. Third: age-appropriate student assent, not assumed from enrolment in the class. Fourth: data anonymisation at the point of analysis - no essay excerpt, score trajectory, or revision plan in any report should be traceable to an individual student. Fifth: disclosure in any publication that the data was generated through a commercial AI service with all that implies for replication and generalisability. These are not bureaucratic formalities. They are the minimum conditions that distinguish classroom research from classroom surveillance. A teacher who shares student score trajectories in a conference presentation without institutional approval and parental consent has used student data without authorisation, regardless of whether the data is anonymised. The ethics obligations exist before the research question, not after.
The dimension that most consistently shows the least improvement between Draft 1 and Draft 2 is Task Achievement - and the reason is architectural rather than motivational. Task Achievement is the dimension that requires the student to revisit the original task prompt and assess whether their essay genuinely addresses it. This is a higher-order judgment that most students do not make spontaneously between drafts. They revise what the feedback agents told them to revise - grammar, vocabulary, structure - and assume the content is adequate because they wrote it. Task Achievement also resists the Revision Plan mechanism in a specific way: a student who did not understand the task on Draft 1 will often not understand it differently on Draft 2 without explicit teacher intervention. The agents can flag that the essay does not address all parts of the prompt, but they cannot explain why the student's interpretation of the prompt was incomplete, or help the student understand what the question was actually asking. This is the strongest argument for the teacher's mandatory monitoring role. A student who scores 1 on Task Achievement on Draft 1 should receive direct teacher contact before the Revision Plan is written, not after Draft 2 is submitted. The three-category monitoring protocol exists precisely for this: a Beginning score on Task Achievement is the signal that requires a human conversation, not just a feedback report.
The claim is ambitious and the evidentiary bar is higher than the chapter's single-course design can meet. Better essays is demonstrable within a course by comparing Draft 1 and Draft 2 rubric scores. Better writers requires evidence that the improvement transfers to contexts where the system is absent - to AI-free timed writing, to writing in other courses, to writing tasks the student encounters beyond the classroom. The evidence that would most convincingly demonstrate the stronger claim would come from three sources. First: a comparison of timed, AI-free writing samples from the beginning and end of the course, scored on the same rubric. Improvement on this measure that parallels improvement on the system's own scores would suggest that what the system is building is genuine competence, not rubric-gaming. Second: student self-assessment data from the Revision Plans across a full course, analysed for increasing sophistication in how students describe their own writing problems and revision decisions. A student who, by Week 12, can name their most persistent error pattern by type, identify its L1 source, and describe the strategy they use to address it has developed metacognitive awareness that transfers. Third: scores on a comparable essay task at the beginning of the following course, assigned before any AI system is used. If students who used the system the previous term outperform those who did not, the transfer claim is supported. This is the longitudinal study the chapter calls for in Section 21.8.
The colleague's position is defensible and should not be dismissed. An institution without a data policy on external AI APIs has not decided that use is permitted - it has not yet decided anything. Acting before a decision exists places the teacher in an exposed position if the eventual policy is prohibitive, and it generates student data under conditions the institution has not authorised. The navigation that preserves both the teacher's professional standing and the students' data protection is: do not wait passively, but do not act unilaterally either. Write to the relevant authority - head of department, data protection officer, IT lead - describing exactly what the system does (student essay text is sent to the Gemini API; only the student ID, not the name, travels with the text; data is retained in Google Sheets for one term), asking for written confirmation that this is permitted under current institutional policy, and proposing the four minimum safeguards from Section 21.7 as the operating protocol. This approach does three things simultaneously: it creates a paper trail that demonstrates the teacher sought authorisation; it typically produces a response within two weeks, which is fast enough to begin mid-term; and it initiates the institutional conversation that will eventually produce the policy, rather than waiting for someone else to start it. If the response is 'wait for the policy,' wait. If the response is 'this is fine under current data handling guidelines,' proceed with documented safeguards. Either outcome is professionally defensible.
This is exactly the interaction the monitoring protocol is designed to generate, and the correct response is not to adjudicate between the student and the agent but to use the disagreement as a teaching moment. First: read the student's thesis statement and the agent's comment together with the student. This takes two minutes and produces one of two outcomes. Either the agent is right and the student does not yet have the metalinguistic vocabulary to recognise the error - in which case the teacher names the error type specifically and connects it to the student's grammar knowledge. Or the agent is wrong - which happens, and the chapter is explicit about the system's accuracy limitations - in which case the teacher confirms the student's judgment, explains why the agent's analysis was incorrect, and uses the moment to build the student's confidence in their own critical reading of AI feedback. Second, regardless of which outcome: record the interaction. A Grammar Agent that consistently misidentifies correct thesis statements as grammatically erroneous has a systematic accuracy problem in that dimension. Three or four such reports from different students in the same cycle is the signal to revise the Grammar Agent's prompt before the next assignment. The student who comes to challenge AI feedback is doing exactly what the system is designed to produce: a writer who reads feedback critically rather than accepting it passively. That student deserves a response that rewards critical engagement, not one that defers to the algorithm.
The protocol is calibrated for a context with the following features: a class of twenty-five to thirty students, one assignment cycle per two to three weeks, a teacher with one to two other classes, and students who are broadly within the same proficiency band. In these conditions, the monitoring burden of the three-category protocol is approximately fifteen to twenty minutes per cycle - manageable and sustainable. Contexts that require a different threshold are common and the teacher should not feel that deviation from the protocol is a failure. A class of forty students doubles the monitoring burden of Category 1 alone. A class with a wide proficiency spread means that the Beginning threshold (score of 8 or below on two or more dimensions) may be too low for an A2 class, where scores of 8–10 represent expected performance, and too high for a B2 class, where any Beginning score is a significant flag. A class that includes students with identified learning needs may require individual monitoring protocols that the three-category system does not accommodate. The principle behind the protocol is more durable than its specific thresholds: every student who cannot benefit from the automated cycle without human intervention must receive that intervention before the revision window closes, not after the final grade is recorded. What counts as 'cannot benefit without intervention' is a professional judgment the teacher makes in light of their specific class. The protocol gives a starting point; the teacher's knowledge of the students gives the calibration.