Thursday, October 8, 2026
Cover illustration for “Debating in English as a Second Language”
Lingual DiscussionDebating in English as a Second Language

Debating in English as a Second Language

Classroom structure leaves ESL learners too little time to develop spoken fluency.

Editorial team · · 11 min read

A classroom of thirty students, one teacher, and a fifty-minute period leaves almost no room for any single learner to speak at length. That arithmetic, not a failure of method or motivation, is the first reason spoken fluency stalls in conventional ESL instruction. Student Talking Time is constrained by class size, by the teacher-fronted structure of most lessons, and by the dominance of written assessment in how progress gets measured. A teacher who lectures, corrects, and tests through written exercises can run an orderly classroom without ever creating the conditions for an individual learner to produce extended, unscripted speech.

Reading and listening are receptive skills. Speaking is productive, and it is trained through an entirely different mechanism. A learner can answer comprehension questions correctly, pass a grammar quiz, and still freeze when you ask them to build a spontaneous argument out loud. Scaria et al.'s Learning in Blocks framework names this directly: quiz-based progression can advance a learner through a curriculum while leaving the gap between recognizing correct language and producing it under real conversational conditions entirely unaddressed. A learner can accumulate vocabulary and grammatical rules for months and still lack the capacity to deploy them in real time, because the exercises that moved them forward never asked them to.

The cost of that gap is rising. Shashidhara et al.'s 2026 field study in Delhi schools found that spoken English proficiency is a powerful driver of economic mobility for low-income youth, connecting directly to job access and wage outcomes. The study also notes a shift already underway: generative AI is lowering the cost of producing polished written English. The written half of language competence is becoming less of a differentiator in the labor market as a result. Spoken fluency is becoming the more valuable and more scarce skill because classroom structure is least equipped to produce it. A system that cannot generate enough Student Talking Time to train it is a structural mismatch between what instruction delivers and what the economy now rewards.

How debate builds fluency

Debate does not simply offer more speaking time. It changes the kind of cognitive work a learner has to do while speaking. Mirzaeva's 2026 synthesis describes structured debate as a catalyst for cognitive restructuring, and the mechanism is specific: constructing a rebuttal in real time forces language processing to move from controlled, monitored production toward automatic fluency faster than other exercises manage. A learner who has to listen to an opponent's claim, identify its weakest point, and respond before the moment passes cannot afford to translate internally or search consciously for grammar rules. That pressure forces a shift that slower, untimed practice only reaches little by little.

That pressure trains three sub-skills simultaneously, which most other formats train one at a time. Argument construction requires selecting and sequencing reasons under a time limit. Refutation requires catching a logical gap in an opponent's position while that opponent is still speaking. Delivery under pressure requires controlling voice, pace, and composure while being actively challenged. A grammar drill isolates form. A scripted dialogue isolates vocabulary retrieval. Debate asks for all three at once, under conditions that do not pause for the learner to catch up.

Moderated group discussion, the format most ESL programs reach for as an alternative to lecture, does not reproduce this pressure. Gao et al.'s study of an ESL conversation club found that learners in group settings often struggle to engage because of language barriers, and that active acknowledgement and encouragement from a moderator was the most effective strategy for improving the quality of discussion. That finding is a genuine success for moderation as a technique, but it reveals the format's limit. The moderator who steps in to encourage a hesitant speaker, rephrase a stuck question, or draw a quiet learner into the conversation is also absorbing the productive demand that the learner would otherwise have to meet alone. Debate removes that crutch by design. No moderator can take a rebuttal for a debater. The opponent's argument still has to be answered by the person it was addressed to, on the clock, without rescue.

The specific speaking gains ESL learners make through regular debate practice

The gains from debate practice do not show up one skill at a time. Mirzaeva's 2026 research synthesis reports that ESL learners engaged in regular competitive and collaborative debate show gains across lexical range, grammatical complexity, and phonological clarity simultaneously, which matters because conventional instruction tends to isolate exactly these dimensions into separate units, separate weeks, and separate exercises.

Lexical range expands under debate conditions because debate topics demand precise vocabulary, and you cannot paraphrase around it. A learner discussing an economic policy cannot substitute a vague gesture for the term they are missing and still win the point. The argument needs the specific word, so that need drives acquisition.

Grammatical complexity increases because rebuttal is built on conditional and causal structures that ordinary conversation rarely calls for. A debater forming a response along the lines of "if they're right that the policy reduces harm, then they still need to explain why the cost falls disproportionately on one group" is producing a sentence shape that a casual exchange about weekend plans will not elicit. Debate makes that structure necessary, not optional.

Phonological clarity improves because an opponent who does not understand a debater's pronunciation will respond to the wrong claim, and that failure is immediate and visible to everyone in the round. Classroom drills simulate correctness without consequence. Adversarial exchange supplies the feedback loop naturally, because being misunderstood costs the argument.

There is also a gain specific to how debate assigns positions. Being forced to argue a side a learner does not personally hold is a linguistic and cognitive forcing function in its own right. It requires active research into unfamiliar material, acquisition of vocabulary the learner would not otherwise reach for, and the construction of an argument without falling back on memorized personal stories, which is the default crutch of most conversational ESL practice.

Why explicit correction undermines speaking confidence

Most ESL classrooms run on explicit correction: a learner speaks, makes an error, and is interrupted with the grammatically correct form. Park et al.'s AI Twin study, presented at CHI '26, draws on prior second language acquisition research suggesting that this interruption discourages risk-taking and slows acquisition, because it reintroduces the anxiety that speaking practice exists to reduce. The study's own data reinforces the point: explicit correction produced significantly lower emotional engagement than rephrasing did.

Krashen's Affective Filter Hypothesis, cited in Park et al., explains why this matters beyond comfort. Elevated anxiety acts as a filter that blocks input from converting into acquisition. A learner who is tense every time they speak is not simply having a worse experience. The tension itself interferes with the process that practice is supposed to produce, so a classroom that corrects aggressively can actively suppress the gains it intends to create.

Rephrasing, often called recasting in the SLA literature, offers an alternative. Instead of halting the exchange to name the error, a more fluent version of the learner's own utterance is modeled back in response, which keeps the conversation moving while still making the correct form audible. The correction happens inside the flow of speech itself.

Park et al.'s AI Twin system tested this with adult ESL learners at CHI '26 by delivering the rephrased utterance in the learner's own voice, a design built to align with what SLA research calls the Ideal L2 Self: the version of the learner who already speaks fluently. That alignment produced higher emotional engagement and higher self-reported motivation than explicit correction did, without interrupting the practice session itself. The lesson extends past one system. If feedback scores an argument instead of halting it to fix the grammar mid-sentence, it can preserve the adversarial pressure debate depends on, while it avoids the anxiety that live correction brings.

How AI judging delivers scored, post-round feedback

The design principle that makes AI judging compatible with everything the previous section establishes is simple to state: the verdict arrives after the final speech, not during it. The round runs uninterrupted from start to finish, the learner's affective filter stays low while they are actually speaking, and the correction, such as it is, arrives only once the adversarial work is done.

The rubric behind that post-round feedback is what makes it useful. Scaria et al.'s Learning in Blocks framework applies CEFR-aligned rubrics to open-ended conversation, scoring Grammar, Vocabulary, and Interactive Communication as independent dimensions, and the resulting judgments agree closely with ESL expert human raters: a degree of variation of 0.23 and a recommendation acceptability rate of 90.91% in Scaria et al.'s analysis. Scoring against a published standard, rather than against a single rater's intuition, is what turns the feedback into something a learner can act on.

Judgment distributed across a panel improves accuracy further than judgment concentrated in one model. Multi-agent debate panels, in which several AI models evaluate a round independently and then reconcile disagreements, outperform single-model judging because the agents challenge each other's conclusions before a consensus score is produced, which approximates the way a panel of human judges works through disagreement before returning a verdict.

None of this amounts to a claim that AI judging replicates human judgment exactly. AI judges still carry positional bias and may not track human judgment reliably in genuinely subjective domains, a limitation the underlying research acknowledges. The response to that limitation is structural: panel design using multiple independent models, rubrics published before the round begins, and a path to human appeal when a learner disputes the result. Trust in the verdict comes from those conditions, along with a platform that has no stake in who wins a given round, not from a claim of parity with a human judge.

Sequencing debate practice from first attempt to live opponent

No solo exercise, however well designed, reproduces the one condition that defines spoken fluency under pressure: an opponent who responds in ways the learner cannot fully anticipate. That unpredictability has to be trained against directly. Practice has to be sequenced.

Solo AI rounds make sense as the first stage. If a learner practices argument construction and delivery without a live opponent in the room, they can build the topic-specific vocabulary and phrase patterns they will need before the stakes rise. That stage should be understood for what it is: preparation, not the destination. A learner who stays there indefinitely is avoiding the unpredictability that live debate exists to train.

Leaderboard and community rounds raise the stakes in a way that sharpens delivery and composure further, because competitive context changes how a learner performs even when the opponent is still mediated through a platform. Reaching a top rank on such a leaderboard can open access to matchups against creators and streamers, which introduces a public-performance dimension that mirrors the pressure of genuinely high-stakes real-world speaking, the kind a job interview or a public presentation demands.

Between sessions, a short self-review habit adds up what you gained from each round. Recording a brief monologue on the round's topic and replaying it with a single focus, pronunciation in one session, grammatical range in another, fluency in a third, turns a written judge's decision into something concrete, because the learner is hearing the same speech the judge actually scored. That habit costs a few minutes and closes the gap between an abstract score and the specific sounds and sentences that produced it.

Choosing debate topics that match CEFR level and maximize productive speaking demand

Topic choice is not incidental to debate practice. ESL educators tier topics by CEFR level because linguistic demand, the vocabulary required, the syntactic complexity a fluent argument needs, the evidence a claim has to marshal, scales with how abstract the topic is. A topic pitched too far above a learner's level produces silence, because there is no available language to build an argument with. A topic pitched too far below it produces fluent, comfortable speech that carries no productive pressure.

At the B1 and B2 levels, topics that require marshalled evidence do the most useful work. Whether AI should be trusted in high-stakes decision-making, or whether social media platforms should enforce stricter age restrictions, are both topics that cannot be answered with a story about the learner's own life. They require conditionals, concessions, and causal chains, exactly the structures that debate practice is positioned to train and that a topic like "describe your favorite holiday" never calls for.

What multilingual AI debate research reveals about argument development

Lai's 2026 study of LLM debates conducted controlled eight-turn debates across six UN languages, covering 71 motions, and found a pattern specific to Chinese: later turns in Chinese-language debates show consistently higher similarity to prior arguments than later turns in English do. Debaters operating in Chinese were more likely to return to an earlier claim in new wording than to introduce genuinely new argumentative ground, and the pattern held across different model agents and topic types.

For an ESL learner whose first language is Chinese, or any language that shares this tendency, the implication is concrete. A rhetorical habit that may be entirely well-formed in the learner's native argumentative tradition, restating a claim in refined language rather than extending it, can read as mere repetition to an English-language adjudicator trained to expect new ground with each turn. That gap is a difference in what counts as argumentative progress from one language tradition to another, and it is learnable once it is named.

That makes the distinction between restating a claim and developing it a specific, trainable skill for ESL debaters, not a vague stylistic preference to be picked up by osmosis. It also sets a concrete requirement for the judging systems built to support these learners: scoring has to track argument development turn by turn, not just the quality of the final speech. A rubric that only looks at the last exchange cannot tell a learner whether they advanced their case or simply repeated it more elegantly, and that distinction is precisely the one this research shows matters most.

Access to live, judged debate practice and school resources

The demand for this kind of practice is not in question. Shashidhara et al.'s field study in low-resource Delhi schools found that students, teachers, and principals alike want tools that support spoken English practice. What the same study found alongside that demand was the structural obstacle: limited devices, intermittent connectivity, and large class sizes make sustained speaking practice nearly impossible to deliver inside the school system alone, no matter how motivated the people inside it are.

Competitive debate has historically compounded that obstacle. Coaching staff, tournament entry fees, and travel budgets are the normal cost of entry into organized debate programs, and that cost concentrates access in well-funded schools. The skill development debate provides, the lexical range, the grammatical complexity, the composure under live challenge, has therefore been distributed by institutional wealth, not by a learner's aptitude or will to practice. A learner in a Delhi classroom with an unreliable internet connection and no debate team to join has the same underlying need as a learner in a well-resourced private school with a full coaching staff. What separates them is access to a live, adversarial, judged round, delivered at a cost and in a format that does not require an institution to stand behind it first.

Sources

  1. USING DEBATES AND DISCUSSIONS TO IMPROVE SPEAKING SKILLS IN ESL CLASSROOMS
  2. AI Twin: Enhancing ESL Speaking Practice through AI Self-Clones of a Better Me
  3. Voice-Based Chatbots for English Speaking Practice in Multilingual Low-Resource Indian Schools: A Multi-Stakeholder Study
  4. Do LLM Debates Repeat Arguments Differently Across Languages?
  5. Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
  6. Moderation Matters:Measuring Conversational Moderation Impact in English as a Second Language Group Discussion
  7. AI Twin: Enhancing ESL Speaking Practice through AI Self-Clones of a Better Me
  8. Voice-Based Chatbots for English Speaking Practice in Multilingual Low-Resource Indian Schools: A Multi-Stakeholder Study

More in Multilingual Debate