The Augmented Educator Podcast

Michael G Wagner

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com

  1. 1d ago

    They Called the Internet a Fad, Too

    Over the past few months, I’ve watched more and more colleagues quietly step back from AI. Not with a dramatic announcement, just with a shrug. The chatbot continues to write mediocre lesson plans. Students use AI to cheat. In general, the hype has cooled. The conclusion feels all too obvious to them. The technology has plateaued, the moment has passed, and their limited professional energy is better spent elsewhere. I understand the impulse. And I fully acknowledge that the numbers seem to support that sentiment. Most organizations have now experimented with generative AI. And yet, a 2026 Gallup survey found that only about one in seven U.S. employees use it daily, and nearly half say they never use it at all. If that’s the reality after three years of relentless promotion, surely the revolution must have been oversold. But this is where the mistake lies. Because that conclusion confuses the interface with the technology. The consumer-facing chatbot, the text box we all learned to use, has indeed largely stopped surprising us. But beneath that familiar interface, the technology is accelerating in ways most of us never get to see. The thing is that we have been here before. Twenty-five years ago, in fact. And the last time we made this mistake, the people who checked out spent the next decade catching up. The winter of 2000, when the internet died On December 5, 2000, at the height of the dot-com collapse, the Daily Mail ran an article declaring the internet a passing fad. This was not just tabloid provocation. The piece summarized findings from the Virtual Society project, an academic study spanning twenty-five European and American universities. Its director, Steve Woolgar, documented widespread user drop-off. Early web surfers had satisfied their curiosity, realized there was more to life offline, and abandoned their modems. His colleagues added that email, far from delivering the paperless office, had mostly delivered information overload. At that time, the market agreed. Between March 2000 and October 2002, the Nasdaq lost roughly 77% of its value. Pets.com, Webvan, WorldCom, and Global Crossing, just to name a few, went under. To a reasonable observer in 2001, the digital economy looked like a collective delusion that had finally been exposed. It was a reasonable interpretation at that moment. But as we know now, this was utterly wrong. What the crash destroyed was the speculative valuation, and not the technology underneath it. The bubble’s real legacy was a massive, debt-financed overbuild of physical infrastructure consisting of millions of miles of fiber-optic cable, much of it laid after the Telecommunications Act of 1996. A lot of that cable then sat unused as “dark fiber” while the companies that laid it went bankrupt. That oversupply then drove the unit cost of bandwidth toward zero, carrying broadband into ordinary homes through the mid-2000s. The effects were measurable. An econometric study of OECD countries, published in The Economic Journal, estimated that each ten-point rise in broadband penetration lifted annual per-capita GDP growth by roughly a percentage point. New technologies, including Web 2.0, e-commerce, streaming, and the cloud, were all built on infrastructure that was financed during the pre-crash mania and dismissed during the post-crash hangover. Public disillusionment had peaked while the real transformation was still being built. It was just out of sight. The people who wrote off the web in 2001 were not wrong about 2001. They were wrong about the decade that followed. The trough looks the same from the inside Generative AI in 2025 and 2026 looks a lot like that moment. S&P Global Market Intelligence reports that the share of companies abandoning most of their AI initiatives more than doubled last year. An MIT initiative, Project NANDA, found that 95% of generative AI pilots failed to deliver measurable financial results. And the RAND Corporation, a nonprofit research organization, puts the failure rate at more than four in five projects, roughly double the rate for conventional IT. The figures all seem to point in the same direction. But if we look a little closer, we start to see a different story. Most failures appear to be organizational rather than failures of capability. Companies bought licenses without deciding what problem they were trying to solve. They deployed models on top of siloed, poorly governed data. And they ran impressive pilots, but assigned no one to own them afterward. A field experiment led by Fabrizio Dell’Acqua showed that while generative AI produces large gains on tasks inside the model’s competence, it can produce negative results on tasks just outside it. This is a boundary the authors call the “jagged frontier.” Failure in AI deployment therefore tells us as much about the organization as it does about the tools being deployed. Meanwhile, the infrastructure story mimics what happened with the early internet. Hyperscalers are projected to spend in access of $700 billion on data centers, GPU clusters, and energy in 2026, and skeptics reasonably ask when the returns will arrive. But just as the fiber surplus collapsed the price of bandwidth, the compute buildout is driving down the cost of running capable models. A workload that cost about $60 per million tokens in late 2021 now runs for cents at comparable performance. And as the unit price fell, usage surged. Aggregate enterprise token use at OpenRouter reportedly grew more than fivefold within six months, to more than 25 trillion tokens per week. 25 trillion tokens per week and rising sharply. That is not what a plateau looks like. The interface plateaued but the technology keeps going So what is all that consumption used for, if not the chatbots we’ve all grown tired of? It is going into agentic systems. Rather than waiting for a prompt, answering, and forgetting, an agent can take a broad objective, break it into steps, use tools across different applications, and keep working toward a result. A conventional chatbot answers one turn at a time. An agent can remember, act, and follow through across many steps without human intervention. You can already see the shift wherever usage data is available. OpenAI reports that agentic workflows overtook conversational usage among its staff within a year, with core departments now generating most of their tokens through agents rather than chat. The shift is also changing internet infrastructure. Cloud providers are launching stateful runtimes for long-running agent workflows, and payment networks are designing “Know Your Agent” frameworks so that autonomous software can transact within regulated payment systems. Peer-reviewed research studies point in the same general direction, though they don’t confirm those exact usage figures, and they usually measure AI-assisted work rather than autonomous agents. Shakked Noy and Whitney Zhang, in a randomized experiment published in Science, found that participants completed professional writing tasks about 40% faster and produced better work while doing it. Erik Brynjolfsson and colleagues measured roughly a 15% increase in issues resolved per hour among customer-support agents. And three other field experiments conducted at software companies, published in Management Science, found developers completing about a quarter more tasks. Taken together, the message is simple. AI is moving from novelty to an everyday tool. But it is doing so in enterprise back offices and developer terminals, well outside the view of an educator who mostly encounters AI through a free chatbot that feels the same as it did last year. The fork in the road to 2030 If you take one idea from this article, it should be this one. The workforce outlook for 2030 describes a fork in the road, depending on whether people adapt as quickly as AI advances. In one scenario, call it the age of displacement, technology outpaces workers. Businesses automate routine cognitive tasks to cut costs. The gains concentrate among a small group of specialists, and everyone else competes for a shrinking pool of narrow, task-based roles. In the second scenario, call it supercharged progress, adaptation keeps pace. Roles are redesigned rather than deleted, and professionals move from doing every task themselves to directing systems of agents. Here is the uncomfortable part. The direction of the economy is beyond people’s control. They can, however, select how they prepare for it. A professional who disconnects now risks preparing only for the displacement scenario. The internet skeptics of 2001 had a plausible excuse for this move, because back then the evidence was still ambiguous. Today, the direction of AI development is clearer, even if the timing is not. But what does preparation actually mean? The research points to a specific set of skills. Work published in Patterns by Zhicheng Lin and colleagues identifies three meta-skills that shape whether AI improves or degrades professional work: setting direction, judging quality, and checking the machine’s output. A related 2026 study finds that such “AI interaction competence” predicts productivity gains better than domain knowledge. You could also call this the agent orchestration skill. That skill doesn’t outsource thinking. It delegates execution to the agents but keeps judgment with the humans. There is one caveat. A 2026 paper introduced the idea of an “augmentation trap.” Early gains encourage adoption, while long-term passive use can erode the expertise that made those gains possible. But orchestration only works when people keep exercising the underlying AI interaction competence. There is no version of this in which you engage once, feel competent, and then just coast. Restricting AI is not the same as ignoring it To be clear, my argument isn’t for implementing AI in every classroom. A teacher may have excellent reasons to restrict AI in a particular course, and

  2. Aug 11

    Why Some AI Drafts Resist Editing

    A few months ago, I published “A Year of AI-Assisted Writing,” a piece describing my AI-assisted writing process. It was a follow-up to my ethics statement, and it laid out in some detail how a blog post moves from an idea in my head to a finished piece on The Augmented Educator: an AI-assisted draft followed by heavy, iterative human editing. I wrote it because many Substack authors appear to work the same way without ever saying so, and the approach still feels controversial enough that I wanted mine disclosed properly. What I did not talk about in that piece is that not every idea that starts in my head makes it onto The Augmented Educator. Sometimes the topic turns out to be less interesting than I thought. Sometimes it ends up a tad too technical. I have a drawer full of essay concepts about cybersecurity in the AI age, but most of them would not appeal to the audience of this Substack. More often than not, however, an idea dies because the initial AI-generated draft resists human editing at a level I did not expect when I started using this workflow. There are drafts that simply fall into place, where the editing feels natural and a few iterations produce a consistent piece that flows. And there are drafts that look polished on the surface but fall apart the moment you start cleaning things up. It is not unheard of for me to reach a point where I simply give up. A point where the editing effort gets me nowhere, where every attempt to fix one problem only surfaces two new ones. People sometimes reach for the saying that “you cannot polish a turd,” and for a long time that was my private shorthand too. But the saying does not quite describe the problem. A turd announces itself. Nobody picks one up expecting to polish it. The drafts I am talking about look clean, read well, and pass every quick inspection, and the trouble only shows once the polishing has begun and hours are already spent. The draft was never bad in any traditional sense. It was just not a starting point from which my iterative workflow had any chance of converging on a piece I would consider fit for my readers. The saying, if anything, gets the situation backwards. The problem is that these drafts look eminently polishable. They just aren’t. I have always wondered why that is, and how to make sure every draft I generate can become a publishable essay rather than a mess of never-ending edits. I suspected the answer might also explain why many professional writers — people who can write perfectly well without assistance — so often struggle with editing AI output. To be clear, I am not suggesting they should trade their tried-and-true approach for an AI workflow. But if there were more clarity about why AI text sometimes resists human editing, it could open up better pathways for teaching AI literacy to professionals who have tried these tools and found them wanting. So in today’s essay, I want to dig into the question of why some AI-generated texts appear polished on the surface yet are fundamentally flawed to the point of being uneditable, while others just work. And I want to explore how to raise the odds that a prompt produces an internally consistent draft in the first place. When fluency is a disguise The first reason an editor might struggle with an AI draft lies in how the text is made. Large language models generate prose by predicting the next token, one after another, in whatever sequence is statistically most likely given the training data. This mechanism is excellent at reproducing the surface of authoritative writing. Computational linguists have a name for the result: deceptive fluency. A deceptively fluent text is grammatically clean, smoothly connected, and formatted exactly as its genre demands, yet hollow underneath. In academic and educational contexts, this produces what some researchers call the fluency fallacy: our tendency to mistake coherent academic language for genuine understanding. Reviewers of AI-drafted literature reviews report the pattern again and again: generic explanations, repetitive sentence structures, weak critical analysis, and conclusions broad enough to fit any paper. The model can summarize ten studies in seconds. It almost never notices the tension between two of them. An editor who sits down to polish a draft is operating on an assumption: that the draft has a sound foundation, and that the remaining work is surface work. Fix the phrasing, add domain expertise, sharpen the examples. With a deceptively fluent draft, that assumption is false. The correct punctuation and the smooth transitions sit on top of an argument made of filler, statistically probable sentences arranged in the shape of reasoning. And the surface is stubborn. Because the model’s prose is so tightly woven at the sentence level, inserting one genuinely analytical thought tends to break the flow of everything around it. You fix a paragraph and the section stops hanging together. You fix the section and the essay’s through-line snaps. The draft gets abandoned, in the end, because its artificial coherence cannot carry the weight of actual reasoning. Untangling the machine’s surface logic costs more than articulating the thought from scratch would have. The inspector on the conveyor belt To understand why tearing down and rebuilding a draft is so exhausting, it helps to look at what post-editing does to the writer’s brain. Traditional models of writing describe a cycle of planning, translating ideas into text, and reviewing. An AI-first workflow reshuffles this cycle. The writer stops being a creator and becomes a reviewer of someone else’s output, and that shift changes the cognitive economics of the whole task. Cognitive load theory sorts mental effort into three kinds: intrinsic load, the inherent difficulty of the task; extraneous load, the wasted effort imposed by bad tools and friction; and germane load, the productive effort that builds understanding. The promise of AI drafting is that it absorbs the intrinsic load of getting ideas into words, freeing the writer’s working memory for higher-order thinking. For some writers, this promise holds. Studies of second-language learners, for instance, find that AI assistance genuinely lifts the burden of grammatical mechanics, and the learners notice it. For an experienced writer editing a full draft, something stranger happens. The typing effort drops, but the effort of evaluation goes up, and it goes up a lot. Reading AI output is not like reading a colleague’s draft. With a colleague, you can trust that there is an intent behind every paragraph, a lived experience, a mental model you share. With a model, you can trust none of that. Every claim might be hallucinated, every transition might be papering over a gap, every confident sentence has to be checked. Researchers developing cognitive load scales for AI-assisted writing have decomposed this into distinct factors — prompt management and critical evaluation — that did not exist in the older models. The practical consequence is a phenomenon that practitioners have taken to calling AI fatigue or review fatigue. Judging whether a generated paragraph matches your intent requires a stream of small verdicts, hundreds of them an hour. Hold the machine’s logic in working memory, compare it against your own knowledge, spot the discrepancy, plan the fix, repeat. Writing from scratch is a proactive state in which the writer builds an arc of coherence at their own pace. Post-editing puts the same writer in the position of a quality inspector on a conveyor belt that never stops. There is a bitter twist at the end of this. Fatigue degrades exactly the faculty the inspector needs most, which is judgment. A worn-down editor starts trusting the machine too readily. The literature calls this automation bias, and it means the drafts most likely to slip through unfixed are the ones that arrived when the editor had nothing left. The forty-percent line The feeling that a draft is unpolishable is not just a mood. It can be measured, and an entire industry has been measuring it for decades. Machine translation post-editing has long needed to know when correcting a machine’s output stops being cheaper than translating from scratch, and its metrics transfer surprisingly well to AI-assisted writing. The workhorse metrics are “Post-Edit Distance” and “Translation Edit Rate.” Both count the minimal operations, the insertions, deletions, substitutions, and shifts, needed to turn a machine draft into the approved final version. The research on these metrics points to a clear threshold. When edits touch roughly 40 percent of a machine-generated text, the effort of post-editing overtakes the effort of writing from scratch. Past that line, the draft is uneconomical to save. Not as a matter of taste. As a matter of arithmetic. Keystroke counts do not tell the entire story, though, because technical effort and cognitive effort are not the same thing. A single semantic flaw might take ten keystrokes to fix and twenty minutes to find. Translation researchers capture this with the pause-to-word ratio. Cognitively demanding output produces clusters of brief pauses in which the editor is reading, re-reading, and deciding how to intervene. A draft can score well on edit distance and still be a cognitive swamp. This is, I think, the empirical shape of the moment I described in the introduction, the moment of giving up. The writer is holding a disjointed machine narrative in working memory while simultaneously trying to plan the coherent structure that should replace it, and that double duty exceeds what working memory can do. Somewhere, consciously or not, the writer runs the numbers and concludes that the honest estimate is past the threshold. The rational move is to stop editing and start over. The first tracks become the rut So far the problems have lived in the draft. The next one lives in us. Cognitive scientists and des

  3. Aug 4

    The Parrot and the Photograph

    If you have spent any time in developer corners of social media over the last few weeks, you will have run into pxpipe. It is a small open-source proxy, published at pxpipe.dev, that sits between AI coding tools like Claude Code and the models behind them. Before a request leaves your machine, pxpipe takes the bulkiest parts of the prompt, the standing instructions, the tool documentation, or the older conversation history, and renders them as PNG images. Instead of sending the model your text, it sends the model a picture of your text. The point of this is money. On real production workloads, pxpipe reports cutting the total bill by 59 to 70 percent. Its demo shows the same coding session costing $42.21 with plain text and $6.06 with images. Same task, same output. The repository collected thousands of GitHub stars within days of going viral, and developers have spent the past weeks arguing about whether this is a clever hack or an accident waiting to happen. For anyone who has not followed the image capabilities of current AI models, this should sound backward. A picture of a page is surely more data than the page. How can it be cheaper to show a machine a photograph of your words than to hand it the words themselves? And, stranger still, why does the machine read the photograph just as well? So in today’s post, I want to unpack image prompting as a cost optimization technique, and then follow the trail somewhere more interesting than a billing statement. Most readers of this blog will never run a proxy or worry about API pricing. Nevertheless, the very fact that this trick is effective shows something significant about the way these systems process information when they read. It is, I will argue, one more crack in the “stochastic parrot” picture of AI, the idea that a language model is nothing more than a very fluent autocomplete. And it opens a door for educators that has little to do with saving money. Why a picture of words can cost less than the words Language models do not read letters or words. They read tokens, small fragments of text, typically a few characters or a short word each, and commercial AI providers bill by the token. A thousand-word document costs you roughly 1,300 tokens every time you send it. If your prompt includes long instructions or an entire stack of reference material, you pay for all of it on every single request. Images are billed differently. An image costs a fixed number of tokens determined by its pixel dimensions, not by what is in it. A page-sized image holding 150 characters and a page-sized image holding 15,000 characters cost exactly the same. Think of the difference between a telegram and a photograph of the telegram. The telegram bills by the word. The photograph costs the same, however many words are on it. That gap is the whole trick. According to pxpipe’s own documentation, about 48,000 characters of standing instructions cost roughly 25,000 tokens as text and roughly 2,700 tokens when rendered as images. Dense material like code, logs, and structured data packs about three characters into each image token, against about one character per text token. To be clear, nothing shady is happening here. Providers price images by area because that is how their vision systems slice them up into a grid of patches. The pricing simply never expected that anyone would send text through the picture channel in disguise. The strange part is not the price The strange part is that the model reads the disguised text fluently. In October 2025, Yanhong Li of the Allen Institute for AI, Zixuan Lan of the University of Chicago, and Jiawei Zhou of Stony Brook University published a paper with the pleasingly blunt title “Text or Pixels? It Takes Half.” They rendered long text inputs as single images, fed them to off-the-shelf multimodal models, and measured what happened. Token counts dropped by roughly half. Accuracy did not drop at all. That held across very different tasks. On a long-context retrieval benchmark, where the model must find one specific fact buried in a mass of text, the image version scored 97 to 99 percent. On news summarization, the image version matched or beat specialized text-compression tools at the same compression rates. And on one large open model, responses even arrived 25 to 45 percent faster because the model had fewer tokens to process. There is a catch, though. Model vision is not OCR, the optical character recognition technology that scanners use to turn a page into exact, character-perfect text. Reading text through the picture channel works at the level of its essential meaning. It works at the level of gist. The consequence is that it can quietly get an exact string wrong: a long ID, a hash, or a precise number. It will not flag the error. It does not know it made one. Keep that in mind. We will need it when we get to the classroom. The parrot was never supposed to do this I have written in a previous essay about the “stochastic parrot” metaphor and its limitations, so I will keep the recap short. The phrase comes from a 2021 paper by Emily M. Bender, Timnit Gebru, and colleagues, and it names the dominant skeptical view of language models: these systems manipulate the form of language with no grip on its meaning. They predict the statistically likely next token, and everything that looks like understanding is an illusion produced by scale. The critique leans on a real philosophical problem, formalized by Stevan Harnad as the symbol-grounding problem and dramatized earlier by John Searle’s Chinese Room. A system that only ever touches symbols, the argument runs, can shuffle them forever without any of them meaning anything. Every piece of this argument is about text. Token in, likely token out, patterns learned from oceans of strings. Now hold that up against what pxpipe does. When a prompt travels through the image channel, the text tokens the parrot supposedly depends on never enter the model at all. Not one character of the original prompt is present in the input. What arrives is a matrix of pixels, patterns of light and dark that happen, to a human eye, to look like writing. Yet the model recovers the instructions from those patterns, follows the logic, and produces the same multi-step work it would have produced from the raw text. The words were left behind at the door. The meaning got in anyway. The technical explanation is that multimodal models translate everything they receive, words and pixels alike, into the same internal representation, a kind of shared space of meaning that researchers call a latent space. A sentence typed as text and the same sentence photographed off a page land in nearly the same spot in that space. Once inside, the model neither knows nor cares which door the meaning came through. Whatever the system is doing, “completing your string” has stopped being an accurate description of it. It is operating on what the string was about. What this does and does not prove Statistics over pixels is still statistics. A skeptic can reply that the model has simply learned pattern-matching across two channels instead of one, which is more impressive but not different in kind. Grounding a word in a photograph of that word is also not grounding it in the real world. The model that reads “apple” off a rendered page has still never held one. And the gist errors cut both ways. Reading by gist looks charmingly human, and it also shows that the system reconstructs content rather than retrieving it exactly, which a determined skeptic can file under sophisticated mimicry. Fair enough, up to a point. Nothing about image prompting settles the deep questions of machine understanding or consciousness. But what was the metaphor actually claiming? A parrot repeats sounds. It holds no representation of what the sounds are about, nothing that would survive if you changed the medium of delivery. A system that pulls the same logical structure out of a character string and out of a photograph of that string demonstrably holds something the parrot lacks: a representation indifferent to the channel it arrived through. Call that a world model, or refuse to. Either way, the metaphor has stopped describing the machine in front of us, and educators who reach for it should be aware that the ground under it has been shrinking for a while. The worksheet and the whiteboard I would guess that almost nobody reading this blog pays per token. If you use AI through a chat subscription, the pricing arbitrage that made pxpipe famous is invisible to you, and you should not install a proxy to save money you are not spending. Even so, two things carry over to the classroom, and it’s the latter that I find genuinely exciting. The first is a piece of AI literacy. Every time you or a student photographs a worksheet, a handwritten draft, or a page of lab data and drops it into a chat, you are doing exactly what pxpipe does: routing text through the picture channel. The model will read it the way it reads those PNGs, fluently at the level of meaning and unreliably at the level of exact strings. A decimal point can drift. A name can change spelling. No warning appears, because the model reads by gist and does not know what it smoothed over. The practical rule is simple enough to teach in five minutes. Use images when you want the machine to understand something; use text when the exact wording or the exact numbers carry the weight; verify either way. The second is that a drawing can now serve as a prompt in its own right. In February 2026, David H. Smith IV and colleagues at Virginia Tech, UC San Diego, and the University of Toronto published a position paper called “Drawing Your Programs.” In a large introductory Python course, students drew problem-decomposition diagrams, boxes, arrows, nested structures, and those hand-built diagrams were fed directly to a model as prompts for code generation. The models handled it well. No translation of the

  4. Jul 28

    Can Kimi K3 Write?

    On July 16, Moonshot AI released Kimi K3, which the company describes as the first open-weight model in the three-trillion-parameter class. As I am writing this, the weights themselves are not yet out. Moonshot has said all 2.8 trillion of them will be published on July 27 under a modified MIT license. I should note that “open weights” does not mean “local execution” in this case. In its native format, the model requires roughly a terabyte and a half of video memory. That is before you count anything else that has to sit in memory alongside the weights. Nobody will be able to run this at home. But what open weights buy is open competition: once the files are public, third-party hosts will be able to serve K3 without asking anyone’s permission, at a launch price low enough that the proprietary labs will have to answer it. For the first time, the open-weight ecosystem is breathing down the necks of Anthropic’s Claude Fable 5 and OpenAI’s ChatGPT 5.6 Sol. That industry story is interesting in itself. But the question I want to address in this newsletter is more personal. Can the thing write? I mean, not benchmark-write. Write under constraints, with a voice a human editor would want to work with. So in today’s post, I want to walk through an experiment I ran to find out exactly that: one research brief, one style guide, three models drafting, four judging blind, and a closing round of guess-who-wrote-what that went almost too well. One brief, no scaffolding Regular readers know the drafting workflow I usually follow. I have described it in an earlier post. I start by sending an essay idea into Gemini 3.5 Deep Research to develop a sourced brief. The brief then goes to a language model, usually the latest Claude model, along with The Augmented Educator style guide. The model returns a rough draft. Only then does the actual writing start. I use an iterative process and rewrite until it sounds like me. The essay idea I used for this experiment came from a YouTube video in which the creator made the following deceptively logical claim: “In art, the effort does not matter. The art itself matters.” His thesis was that talented artists will thrive with AI tools, whereas untalented artists will fall behind regardless of AI use. I felt there might be some historical context to unpack here, since this is likely not the first time this claim was made. It therefore seemed like a good seed for an essay on what happens to art education when machines absorb the effort. Perfect for The Augmented Educator. For the purpose of this experiment, I applied one major deviation from my routine. Usually, the prompt I use for drafting includes detailed instructions about story angle and structure. Here, I withheld all of it. The models got the brief and the guide, nothing more. I wanted to see and evaluate their raw judgment and writing skills, and not my own scaffolding mirrored back at me. The result was three drafts by three contenders: * Claude Fable 5 (run on its Max setting) wrote “Take the Hand from the Picture.” * ChatGPT 5.6 Sol (run on its Pro setting) wrote “The Art Does Not Come With a Timesheet.” * Kimi K3 (run on its Max setting) wrote “The Panel with Three Lines.” The essays themselves were not really remarkable, and to be honest, none of them would make it onto this Substack without very heavy editing, if at all. I am attaching them here only for reference in case someone wants to check them out. But regardless, I felt they were good enough for the experiment. A three-to-one landslide I handed the three unlabeled essays, plus the style guide, to four AI judges for a blind review. These judges were fresh instances of the three author models and Gemini 3.5 Pro, the model that generated the brief. Each judge scored every essay from 1 to 10 across three categories: style-guide adherence, prose quality, and argument and storytelling. Each model was also asked to pick exactly one essay to publish. In the following table, each figure is one judge’s three category scores averaged into a single mark out of 10 for that essay. And the bottom row averages those marks across all four judges. If you look at these numbers, one question pops up immediately. Why did Fable’s draft win so clearly? Two main reasons came up in three of the four verdicts. The first was the model’s sheer discipline. It hit the guide’s fussiest targets, and it hit them visibly: wherever a list wanted three items, Fable wrote four, dodging the banned rule of three. It was also the only draft that dug past the obvious material in the brief and used the specialist studies buried further down, which the other two left untouched. The second reason was intellectual. Fable’s was the only draft that made the logical collapse of the YouTuber’s quote its thesis rather than a passing correction. Because if talent is a fixed quantity that tools merely expose, as the claim indicates, then teaching art is a pointless exercise. Fable noticed this inconsistency in the brief’s analysis and built its entire narrative around it. Sol dissented on both counts. It was the only judge that marked Fable’s draft down on adherence, faulting the long paragraphs and a stack of balanced contrasts. And it thought its own draft handled the talent question more cleanly. The other texts split the judges. Sol’s “Timesheet” drew genuine praise for separating what is worth encountering as art from what is worth assigning as education. But its bolded imperative takeaways read to Fable like a faculty memo, and to Kimi like a workshop handout drifting toward the generic “5 ways to...” article the guide warns against. Fable also found its clipped, uniform rhythm the most machine-like in the pool. Kimi’s “Panel” split the room differently. Every judge praised the liveliness of its sentences. Two ranked the draft second on points, but none of them recommended publishing it first. And Sol liked the individual sentences but disliked the voice they added up to, calling it prosecutorial rather than provocative and short on generosity toward students. Before I continue, I need to add a quick note on bias, because some readers are probably already typing in the comment section. Language models grading language models is somewhat of a circular exercise, and self-preference is a documented failure mode. Sure enough, the two proprietary author models, Fable and Sol, each picked their own work. Gemini, with no skin in the game, sided firmly with the majority. Guess who wrote what After scoring, I asked each judge to guess which essay was written by which model. Three of the four guessed perfectly. Each model, it turns out, wrote with some recognizable habits: * Fable followed instructions to the letter, down to the intentional four-item lists. * Sol leaned on a rigid structure, characterized by short, symmetrical sentences and takeaways formatted as bolded lists. * Kimi K3 wrote with a casual punch and put momentum above the fine print. This is exactly how the banned rule of three slipped back into its essay. Kimi identified “Panel” as its own work because it “treats the guide as a vibe rather than a spec.” That is a sharp piece of self-recognition from a model that had ranked that same essay dead last a few minutes earlier. Only Gemini stumbled. It described the habits accurately but filed two of them under the wrong names, pinning Sol’s bolded lists on Kimi and Kimi’s rule-of-three slips on Sol. The judges also proactively disclosed the limits of the exercise. Sol footnoted an arXiv paper, warning me that model attribution remains an unsolved research problem. Kimi insisted upfront that models have no privileged ability to recognize their own text. And Fable volunteered a disclosure that, as a Claude model, it might be flattering its own family. So can Kimi actually write? Now to the core question. Can Kimi K3 write at the level of the leading flagship models? Let’s start with the good news, which is the voice. Two judges called Kimi’s hook the strongest of the three, and its paragraphs move better than anything else in the pool. Nothing in it sounds like the beige, press-release text we so often associate with machine drafts. As for the bad news, Kimi’s draft broke the guide’s most explicit ban, and broke it repeatedly. The rule of three is back on nearly every page, although that would be easy to fix in post-editing. The larger problem was the logic. The essay endorsed the YouTuber’s claim, then pivoted to a lesson plan anyway. That is the contradiction Fable’s draft made its thesis, and the one Sol’s draft avoided by refusing the fixed-talent premise. Kimi carried it to the final paragraph without noticing. As I have written before, I treat a model draft like a delivery of clay. The material is real, but the shape isn’t mine yet. That clay-carving stage is where the actual writing happens for me, and it is why Kimi’s logical lapses are more consequential to me than its stray triads. What educators should take away This admittedly imperfect experiment suggests that Kimi K3 can write close to the level of the leading foundational models. Its draft finished essentially level with Sol’s and well behind Fable’s, which is still a startling place for an open-weight model to land. However, I will probably stay with Claude Fable 5 for my drafting workflow, at least for as long as Fable is included in my Claude subscription. That could change, though, because open weights alter the financial arithmetic underneath the whole comparison. Fable on its Max setting costs real money. By contrast, Moonshot launched K3 at three dollars per million input tokens and fifteen per million output. And once the files are public, no single vendor controls that meter. It is cheaper, swappable, inspectable, and impossible to un-release. For an educator or a small publication choosing a drafting engine on a budget, K3 is the first open mod

About

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com