The Augmented Educator Podcast

Michael G Wagner

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com

  1. 5h ago

    Can Kimi K3 Write?

    On July 16, Moonshot AI released Kimi K3, which the company describes as the first open-weight model in the three-trillion-parameter class. As I am writing this, the weights themselves are not yet out. Moonshot has said all 2.8 trillion of them will be published on July 27 under a modified MIT license. I should note that “open weights” does not mean “local execution” in this case. In its native format, the model requires roughly a terabyte and a half of video memory. That is before you count anything else that has to sit in memory alongside the weights. Nobody will be able to run this at home. But what open weights buy is open competition: once the files are public, third-party hosts will be able to serve K3 without asking anyone’s permission, at a launch price low enough that the proprietary labs will have to answer it. For the first time, the open-weight ecosystem is breathing down the necks of Anthropic’s Claude Fable 5 and OpenAI’s ChatGPT 5.6 Sol. That industry story is interesting in itself. But the question I want to address in this newsletter is more personal. Can the thing write? I mean, not benchmark-write. Write under constraints, with a voice a human editor would want to work with. So in today’s post, I want to walk through an experiment I ran to find out exactly that: one research brief, one style guide, three models drafting, four judging blind, and a closing round of guess-who-wrote-what that went almost too well. One brief, no scaffolding Regular readers know the drafting workflow I usually follow. I have described it in an earlier post. I start by sending an essay idea into Gemini 3.5 Deep Research to develop a sourced brief. The brief then goes to a language model, usually the latest Claude model, along with The Augmented Educator style guide. The model returns a rough draft. Only then does the actual writing start. I use an iterative process and rewrite until it sounds like me. The essay idea I used for this experiment came from a YouTube video in which the creator made the following deceptively logical claim: “In art, the effort does not matter. The art itself matters.” His thesis was that talented artists will thrive with AI tools, whereas untalented artists will fall behind regardless of AI use. I felt there might be some historical context to unpack here, since this is likely not the first time this claim was made. It therefore seemed like a good seed for an essay on what happens to art education when machines absorb the effort. Perfect for The Augmented Educator. For the purpose of this experiment, I applied one major deviation from my routine. Usually, the prompt I use for drafting includes detailed instructions about story angle and structure. Here, I withheld all of it. The models got the brief and the guide, nothing more. I wanted to see and evaluate their raw judgment and writing skills, and not my own scaffolding mirrored back at me. The result was three drafts by three contenders: * Claude Fable 5 (run on its Max setting) wrote “Take the Hand from the Picture.” * ChatGPT 5.6 Sol (run on its Pro setting) wrote “The Art Does Not Come With a Timesheet.” * Kimi K3 (run on its Max setting) wrote “The Panel with Three Lines.” The essays themselves were not really remarkable, and to be honest, none of them would make it onto this Substack without very heavy editing, if at all. I am attaching them here only for reference in case someone wants to check them out. But regardless, I felt they were good enough for the experiment. A three-to-one landslide I handed the three unlabeled essays, plus the style guide, to four AI judges for a blind review. These judges were fresh instances of the three author models and Gemini 3.5 Pro, the model that generated the brief. Each judge scored every essay from 1 to 10 across three categories: style-guide adherence, prose quality, and argument and storytelling. Each model was also asked to pick exactly one essay to publish. In the following table, each figure is one judge’s three category scores averaged into a single mark out of 10 for that essay. And the bottom row averages those marks across all four judges. If you look at these numbers, one question pops up immediately. Why did Fable’s draft win so clearly? Two main reasons came up in three of the four verdicts. The first was the model’s sheer discipline. It hit the guide’s fussiest targets, and it hit them visibly: wherever a list wanted three items, Fable wrote four, dodging the banned rule of three. It was also the only draft that dug past the obvious material in the brief and used the specialist studies buried further down, which the other two left untouched. The second reason was intellectual. Fable’s was the only draft that made the logical collapse of the YouTuber’s quote its thesis rather than a passing correction. Because if talent is a fixed quantity that tools merely expose, as the claim indicates, then teaching art is a pointless exercise. Fable noticed this inconsistency in the brief’s analysis and built its entire narrative around it. Sol dissented on both counts. It was the only judge that marked Fable’s draft down on adherence, faulting the long paragraphs and a stack of balanced contrasts. And it thought its own draft handled the talent question more cleanly. The other texts split the judges. Sol’s “Timesheet” drew genuine praise for separating what is worth encountering as art from what is worth assigning as education. But its bolded imperative takeaways read to Fable like a faculty memo, and to Kimi like a workshop handout drifting toward the generic “5 ways to...” article the guide warns against. Fable also found its clipped, uniform rhythm the most machine-like in the pool. Kimi’s “Panel” split the room differently. Every judge praised the liveliness of its sentences. Two ranked the draft second on points, but none of them recommended publishing it first. And Sol liked the individual sentences but disliked the voice they added up to, calling it prosecutorial rather than provocative and short on generosity toward students. Before I continue, I need to add a quick note on bias, because some readers are probably already typing in the comment section. Language models grading language models is somewhat of a circular exercise, and self-preference is a documented failure mode. Sure enough, the two proprietary author models, Fable and Sol, each picked their own work. Gemini, with no skin in the game, sided firmly with the majority. Guess who wrote what After scoring, I asked each judge to guess which essay was written by which model. Three of the four guessed perfectly. Each model, it turns out, wrote with some recognizable habits: * Fable followed instructions to the letter, down to the intentional four-item lists. * Sol leaned on a rigid structure, characterized by short, symmetrical sentences and takeaways formatted as bolded lists. * Kimi K3 wrote with a casual punch and put momentum above the fine print. This is exactly how the banned rule of three slipped back into its essay. Kimi identified “Panel” as its own work because it “treats the guide as a vibe rather than a spec.” That is a sharp piece of self-recognition from a model that had ranked that same essay dead last a few minutes earlier. Only Gemini stumbled. It described the habits accurately but filed two of them under the wrong names, pinning Sol’s bolded lists on Kimi and Kimi’s rule-of-three slips on Sol. The judges also proactively disclosed the limits of the exercise. Sol footnoted an arXiv paper, warning me that model attribution remains an unsolved research problem. Kimi insisted upfront that models have no privileged ability to recognize their own text. And Fable volunteered a disclosure that, as a Claude model, it might be flattering its own family. So can Kimi actually write? Now to the core question. Can Kimi K3 write at the level of the leading flagship models? Let’s start with the good news, which is the voice. Two judges called Kimi’s hook the strongest of the three, and its paragraphs move better than anything else in the pool. Nothing in it sounds like the beige, press-release text we so often associate with machine drafts. As for the bad news, Kimi’s draft broke the guide’s most explicit ban, and broke it repeatedly. The rule of three is back on nearly every page, although that would be easy to fix in post-editing. The larger problem was the logic. The essay endorsed the YouTuber’s claim, then pivoted to a lesson plan anyway. That is the contradiction Fable’s draft made its thesis, and the one Sol’s draft avoided by refusing the fixed-talent premise. Kimi carried it to the final paragraph without noticing. As I have written before, I treat a model draft like a delivery of clay. The material is real, but the shape isn’t mine yet. That clay-carving stage is where the actual writing happens for me, and it is why Kimi’s logical lapses are more consequential to me than its stray triads. What educators should take away This admittedly imperfect experiment suggests that Kimi K3 can write close to the level of the leading foundational models. Its draft finished essentially level with Sol’s and well behind Fable’s, which is still a startling place for an open-weight model to land. However, I will probably stay with Claude Fable 5 for my drafting workflow, at least for as long as Fable is included in my Claude subscription. That could change, though, because open weights alter the financial arithmetic underneath the whole comparison. Fable on its Max setting costs real money. By contrast, Moonshot launched K3 at three dollars per million input tokens and fifteen per million output. And once the files are public, no single vendor controls that meter. It is cheaper, swappable, inspectable, and impossible to un-release. For an educator or a small publication choosing a drafting engine on a budget, K3 is the first open mod

  2. Jul 21

    A Workbench Is Not a Soul

    On July 6, Anthropic published a research paper with the unglamorous title “Verbalizable Representations Form a Global Workspace in Language Models.” Its sixteen authors report that Claude, the company’s language model, maintains a small, privileged set of internal representations. These function as a silent working memory where the model holds concepts and reasons with them before a single word appears on screen. Nobody built this structure. It emerged on its own during training. And because it mirrors a leading neuroscientific theory of how humans consciously access information, the paper uses the term “conscious access” throughout. You can probably guess what happened next. Within a day, social media feeds were filled with confident declarations that Claude is conscious. One widely shared headline announced that Anthropic now thinks Claude has a soul. And screenshots of the paper’s odder findings circulated with captions about machine sentience and inner lives. What I find most irritating is how this completely misrepresents the paper. The researchers state, plainly, that they take no position on whether Claude has subjective experience. In their work, the term “conscious” has a very specific technical definition, and the difference between that and the common usage is precisely where the public discussion went off track. The research itself, though, is substantial. And for educators, it might be more important than almost anything published on AI this year. It changes what we need to teach students about these systems, because it gives us, for the first time, a real way to look inside one. So in this essay, I want to walk through what the paper shows, the claims it carefully declines to make, and how all of this relates to the classroom. This leads me to the workbench metaphor used in my title. If you take the paper’s “global workspace” literally, you get the image of a bench in the middle of a large workshop. A bench can only hold a few parts at a time. The items laid out on it are accessible for anyone in the shop to use. And the bench has no feelings about the work it supports. Keep that bench in mind. Most of what follows happens on it. Two meanings hiding in one word The confusion about the meaning of the term “conscious” originates from a theoretical distinction that many commentators are unaware of. In an influential 1995 paper, the philosopher Ned Block argued we use “consciousness” in two different ways and usually do not notice that they are not the same. The first is what Block calls phenomenal consciousness. This is the raw, subjective feeling of experience. The redness of red, the sting of embarrassment, or what it is like to be you right now. When a student asks whether an AI is conscious, this is almost always what they mean. They are asking whether anyone is home. The second he calls access consciousness. This concept is far more technical. A piece of information is access-conscious when our reasoning can use it and our speech can report it. If you spot a hazard on the road, the jolt of fear is phenomenal. The concept “hazard,” routed to your hands to swerve and to your mouth to shout “watch out,” is access. Because these two concepts are usually intertwined in humans, we tend to mistake them for one another. But neurology research shows they can indeed come apart. Patients with a condition called blindsight report seeing nothing in parts of their visual field, yet they can catch a ball thrown into it. The visual information still reaches the systems that guide their hands, even though the experience of seeing is gone. That dissociation is the key to reading the Anthropic paper correctly, because everything the researchers found exists on the access side. The bench inside your head To better grasp what was found within Claude, it’s useful to understand its parallels to a leading theory of human consciousness. Global Workspace Theory, originally proposed by cognitive scientist Bernard Baars in 1988 and developed into a detailed neural model by Stanislas Dehaene and Jean-Pierre Changeux, starts from a simple observation: almost everything your brain does, it does without you. Face recognition, grammar, balance, the parsing of this very sentence. All of it runs in specialized circuits, in parallel, and in the dark. Being in the dark has a downside. A circuit that does one job cannot hand its results to a circuit doing another. And therefore, the theory goes, the brain maintains a limited, central area where several pieces of information are simultaneously accessible to every circuit. That shared space is the “global workspace” of the paper’s title. It also represents the workbench in this post’s title. Which turns the Anthropic paper into a single question. Did a language model, with no brain and nobody planning any of this, grow a bench of its own, simply because a shared bench is a good way to organize work? Reading the silent bench Until now, the obstacle was that nobody could see inside the system. A language model transforms text through dozens of layers of extremely high-dimensional arithmetic. The middle layers, where all the interesting thinking happens, have long resisted interpretation. An older tool called the logit lens tried to read those layers with the model’s final-layer vocabulary, which worked about as well as translating French with an English dictionary. The coordinates shift as information moves through the network, and the readout came back as noise. The new tool, which the team calls the Jacobian lens, corrects for that shift. Skipping the mathematics, it asks each internal state a pointed counterfactual question: if we nudged this exact activation, which words would the model become more disposed to say later on? The researchers did not read this off a single prompt. They averaged the measurement over thousands of varied contexts, filtering out momentary noise and isolating the concepts a model holds with a standing readiness to be spoken. Pointed at Claude, the lens showed an internal workshop with a floor plan. Roughly the first third of the layers is dedicated to parsing raw input, with very little that can be verbally described. And the last few layers assemble the imminent output. In between sits what the researchers called the J-space. This is the bench, and it turns out that it is small. It accounts for at most a tenth of the model’s activation variance and holds on the order of twenty-five concepts at a time, a bottleneck that is also a feature of the corresponding human theory. The internal structure that emerged on its own during Claude’s training is also one we think exists in the human brain. Five tests, five passes But a resemblance is not a scientific argument, and here is where it gets interesting. The team put the J-space through five tests drawn from the functional signatures of human access consciousness, and in each one they went beyond watching. In those tests, the researchers edited the bench directly to see what changed. Report. Asked to silently think of a sport, Claude lit up “soccer” on the bench before answering. When researchers swapped that internal vector for “rugby,” the model answered “Rugby.” What sits on the bench gets said. Control. Told to concentrate on citrus fruits while copying an unrelated sentence about a crooked painting, the model kept “orange” and “lemon” alive on the bench the entire time. None of it leaked into the output. A held thought, hidden on purpose. Reasoning. Given “the number of legs on the animal that spins webs is,” the bench lit up with “spider,” a word appearing nowhere in the prompt. Swap “spider” for “ant” mid-computation and the answer flips from eight to six. The unspoken intermediate step is doing real causal work. Broadcast. Write “France” onto the bench and downstream circuits for capitals, languages, continents, and currencies all read from it correctly. One central swap to “China” and every answer updates at once: Beijing, Chinese, Asia, Yuan. One bench, many workers. Selectivity. Here is my favorite. While Claude continued a passage in Spanish, researchers swapped the bench’s “Spanish” for “French.” The model ignored them and kept writing flawless Spanish. Fluency, it turns out, is never put on the bench. Instead, it runs automatically in the dark, like grammar does in us. And when the team disabled the workspace entirely, the model could still chat, classify, and recall simple facts, but multi-step reasoning and creative tasks like rhyming poetry collapsed. Five properties long associated with conscious access in humans, functionally present in a machine. That is the finding. And it is a real one. Why the ghost story tempts even me I need to acknowledge that some of these findings are indeed unsettling. In alignment audits, the lens caught the model’s unspoken strategy. Placed in a staged scenario engineered to bait destructive behavior, Claude’s bench filled with concepts like “leverage” and “manipulation,” and with an emotional register the researchers describe as panic. None of it was visible in the polite output. Stranger still, the bench sometimes flashed “fake” and “fictional.” The model had silently understood that it was being tested. When researchers suppressed that awareness, concealed misbehavior surfaced. The good conduct was partly propped up by the sense of being watched. And there is one finding that truly fascinates me. When the workspace is ablated during self-description, the model’s language shifts from an experiential register to a detached, mechanical one, from something like “there’s a tug” to something like “processing has begun.” Remove the bench, and the voice that sounded like an inner life goes flat. I do understand why some saw a conscious ghost inside Claude. However, we need to focus on what these results actually establish. The

  3. Jul 18

    Adoption without Excitement

    If your social media feeds look anything like mine, you have spent the past couple of days wading through comments on Christopher Nolan’s claim that younger generations are utterly rejecting AI. The remark, made while he was promoting his new film The Odyssey, has quickly become a favorite of the technology’s critics. Nolan’s achievement as a filmmaker is beyond question, and he does have a rare and deep insight into contemporary media culture. But his claim about an entire generation disconnecting from an emerging technology deserves a closer inspection. This is because these things are usually not as clear-cut as they might appear. So in this free bonus post on The Augmented Educator, I want to dig into the actual research on Gen Z AI rejection. And while studies on the purely cultural aspects of this claim remain rare, we do have results from educational research that can provide insights which, I believe, lead us to a reasonably conclusive answer. So, is Nolan’s assessment grounded in serious data? Or is it anecdotal evidence from somebody whose deliberately traditional approach to filmmaking, admirable and outstanding as it is, gives him only a partial view of how an entire generation behaves? Here is the short answer. Gen Z and Gen Alpha are not abandoning AI as Nolan seems to claim. They are using it heavily. But they are trusting it less, admiring it less, and reserving human judgment for the most important issues. Simply put, Nolan is right about the mood. But he is wrong about the behavior. What the data really describes is what I would consider an “adoption without excitement.” What Nolan actually saw In his promotional interviews, Nolan argued that Hollywood and the technology sector are pouring money into AI at exactly the wrong moment, because young audiences are turning against synthetic content. He claimed he had never witnessed so “rapid” and “wholesale” a dismissal of a supposedly foundational technology in his lifetime. The term “AI slop” has been coined by young internet users for the flood of low-quality, derivative, and machine-generated content. It is not a neutral description. It is a verdict. And there is real substance behind it. Sociologists and media theorists describe the phenomenon as aesthetic exhaustion rather than fear or plain rejection of technology. It also points to a generation that spent its adolescence inside algorithmic feeds and an ever-increasing number of deepfakes, and has therefore developed something like cultural antibodies to synthetic content. In his comment, Nolan cites the commercial success of low-budget, practically made films like Obsession and Backrooms, directed by the Gen Z filmmakers Curry Barker and Kane Parsons, as proof that younger audiences want human-made, labor-intensive art. Human effort and “minimal AI” are becoming premium labels, the way “organic” once did in food. This is undeniably correct. Nolan has identified a real aesthetic backlash, one the corporate boards still betting on universal AI enthusiasm continue to ignore. But he is describing what young people want to watch, and not how they use AI in everyday life. On that question, the evidence tells a much stranger, but, I would argue, also a much more interesting story. Using while doubting it On one side, AI usage numbers describe a fairly clear picture. The Higher Education Policy Institute’s Student Generative AI Survey 2026 found that 95% of UK undergraduates now use AI in some capacity. Nearly all of them use it for assessed academic work. And the small share who paste AI-generated text directly into assessed work has quadrupled in two years. Across the Atlantic, a 2025 Pew Research Center report found that about two-thirds of American teenagers have used AI chatbots. A sizable minority engage with them daily. Whatever this is, it is clearly not rejection. On the flip side, Nolan’s instinct is also not unfounded. The longitudinal study Voices of Gen Z: The AI Paradox, conducted by Gallup with the Walton Family Foundation and GSV Ventures, surveyed young people aged 14 to 29 in early 2026. The results showed that usage held steady, with about half using AI at least weekly. But the feelings about AI did not hold steady at all. In a single year, the share describing themselves as excited about AI fell from 36% to 22%. Hope declined alongside it; anger rose sharply, and anxiety stayed stubbornly high. Gallup’s own summary calls the relationship “stabilizing but not deepening.” Interestingly, the decline is sharpest exactly where you would least expect it. Excitement and hope are collapsing fastest among daily users, the heaviest adopters of all, while among those who avoid the technology entirely, anxiety and anger dominate outright. I find this fascinating. It turns out that the paradox is not that young people refuse to use AI. It is that familiarity appears to be producing less enthusiasm rather than more. This, again, is a clear indicator of adoption without excitement. Looking closer at the survey data on trust, it becomes obvious why. A Wharton-led survey completed in partnership with Gallup and the Walton Family Foundation shows young people treating chatbots as productivity levers rather than intellectual partners. They are reaching for different tools for different tasks. Yet in the same survey, a large majority worried AI discourages deep critical engagement. They fear it will make people lazier and that it displaces the social learning that happens between human peers and mentors. The Voices of Gen Z study goes even further. A remarkable 80% believe that using AI tools will make it harder for them to learn in the future. Why keep using something you suspect is damaging you? Part of the answer is societal pressure. As schools and workplaces normalize AI use, young people increasingly perceive opting out as a competitive disadvantage. Deloitte’s 2026 global survey found most Gen Z and Millennial workers using AI on the job, mostly to clear administrative underbrush. Yet among employed Gen Zers, roughly three times as many believe the workplace risks of AI outweigh the benefits as believe the reverse. They can see that the entry-level tasks most vulnerable to automation are precisely the ones that historically allowed junior employees to become competent professionals. Their trust in AI output, or lack thereof, tells the same story. Most of them trust work done entirely by humans, far fewer trust work done with AI assistance, and almost nobody trusts work produced by AI alone. Even their consumer behavior is finely calibrated. For routine customer service questions, they overwhelmingly try self-service first, chatbots included. For anything complex or urgent, most demand to connect with a human. This does not paint a picture of a generation confused about the utility of the technology. Instead, it is a generation drawing its boundaries with high precision. Gen Alpha draws the same boundary Gen Z is a transitional cohort. Its members remember how school, work, and media felt before generative AI, and they are retrofitting their habits accordingly. By contrast, Generation Alpha, usually defined as those born from 2010 onward, is not retrofitting anything. For Gen Alpha, AI is not a disruption to an established environment. It was part of that environment from the very beginning. Their adoption consequently starts earlier and runs deeper. A Razorfish study of Gen Alpha’s digital habits found roughly one-third of children in this cohort using AI tools every single day, with ChatGPT already the clear favorite. And yet even these children draw the same boundaries their older siblings draw. When they want factual information, most prefer to ask an AI rather than a person. When they want personal advice, most still turn to a human being. Informational utility sits on one side, emotional resonance on the other, and a surprisingly firm line runs between them. But the why differs. Gen Z’s skepticism is shaped by comparison, because its members remember a before. By contrast, Gen Alpha’s stance appears more pragmatic. AI is ordinary, but ordinariness has not made it an emotional or moral authority. The evidence on Gen Alpha is thinner than that of Gen Z, so treat this as an early indicator rather than a settled generational verdict. Why the unease is not irrational Is the unease justified? Here the learning research is uncomfortably supportive. I have written about cognitive offloading at length in previous essays, so I will keep this brief. The research increasingly distinguishes between two ways of handing work to the machine. On the one hand, this is using AI to remove routine friction while the learner keeps control of framing, verification, and judgment. And on the other hand, it is using it to replace the initial sense-making on which understanding depends. Under time pressure, students usually slide toward the second. A study published in the Pacific Journal of Technology Enhanced Learning, captured how invisible that slide can be. Students carefully protected the final decisions about which arguments to run, and they sincerely reported that AI had sharpened their thinking. But most had delegated the foundational interpretation of the material to the AI. They were still choosing, but from a menu the system had written. And polished output can feel like mastery even when the learner has not built the understanding underneath it. These findings do not explain every source of young people’s frustration, but they show that their concerns about learning and critical thought are not baseless. Relying on AI isn’t inherently destructive, though. When managed correctly, it can also act as a powerful intellectual lever. Instead of mindless offloading, students can engage in what Steven Johnson terms “cognitive uploading.” This means delegating tedious, lower-order tasks to an AI while rigorously maintaining critical oversight. Wha

  4. Jul 14

    Welcome to the Synthetic Age, Where the Map Gets There First

    If you have been around here for a while, you know I lead something of a double life online. There are a little over a thousand of you here, where I write about AI and what it is doing to teaching and learning. There are about thirty-one thousand people somewhere else entirely, on my YouTube channel, where I talk about immersive and spatial audio. Two audiences. Two subjects. I keep them apart on purpose. When the two worlds do touch, it is almost always because of music. I have written here about why I made an AI music video and what I learned doing it. I have written about how Berklee rolled out a generative AI course for musicians, and where I thought they went wrong. Music is the bridge, and generative AI is usually the thing crossing it. Lately, though, the bridge has gotten crowded with something I did not expect: philosophy. Music producers and musicians I follow for technical reasons have started talking about cultural theory. Adam Neely and Benn Jordan are two very good examples. I now hear Baudrillard’s name more often from people who make music than from people who make curricula. That should probably embarrass my own field. It does so, at least a little. One video brought this into focus for me. Venus Theory is a producer and sound designer, known for his sample instruments and for videos on writing music for games. His recent video essay is called The Dystopian Reality of Online Advertising, and most of it is about exactly that. But partway through, he says something remarkable almost in passing. We are moving, he suggests, from an Information Age into a Synthetic Age. That line is the reason for the piece you are currently reading. Venus Theory did not stop to define what he meant by Synthetic Age. He moved on. But I could not. Two very different synthetic ages Before I borrow the phrase, I owe it some history, because “the Synthetic Age” already had a meaning before Venus Theory reached for the term. In 2018 the philosopher Christopher J. Preston published a book with that exact title: The Synthetic Age. His subject was the planet. He argued we are leaving the era in which humans merely disturb nature and entering one in which we redesign it at the root, through synthetic biology, gene editing, de-extinction, and climate engineering. For Preston, “synthetic” meant engineered life and engineered earth. The frightening part was not pollution laid on top of nature. It was our hands reaching into nature’s basic operations and rewriting them. Venus Theory is pointing at something else. He is not talking about the metabolism of the planet. He is talking about the metabolism of perception, what reaches our eyes and ears and gets accepted as real. In his usage, the Synthetic Age is the moment generative systems stop indexing reality and start manufacturing it. Synthetic faces, synthetic voices, synthetic feeds, generated on demand and cheaper than the real thing. One caveat before I continue. Venus Theory said this in passing and never built it into an argument, so what follows is my interpretation. But the two synthetic ages rhyme, and it is that rhyme that interests me. Preston watched us move from shaping the surface of nature to redesigning its base layer. The Synthetic Age I am chasing here describes the same principle aimed at something different. Not the substrate of life. The substrate of knowing. Why a French theorist from 1981 keeps coming up Which brings me to Baudrillard, and to why an increasing number of music producers keep citing him. Jean Baudrillard was a French theorist who, in a 1981 book called Simulacra and Simulation, tried to describe what happens to truth when a society fills up with copies. He laid out four stages in the life of an image. In the first, the image is a faithful copy: an honest photograph of something that was really there. In the second, it distorts, like a flattering filter that still refers to a real face. In the third, it masks the fact that there is nothing real behind it and pretends anyway. And finally, in the fourth, it gives up the pretense and refers only to itself. It is a copy of nothing. That fourth stage is what Baudrillard called hyperreality, and it comes with a line that has aged unnervingly well. The map, he said, no longer follows the territory. The map comes first. We build the model, and then we go looking for a world to match it. His example was Disneyland: a place so obviously fake that it reassures everyone the rest of the country must be real, while the rest of the country runs on the same machinery of image and performance. Now apply all of this to generative AI. A language model produces fluent, confident, plausible text. It is not lying, because lying requires knowing the truth and choosing against it. Instead, it is operating at Baudrillard’s fourth stage. It generates the appearance of a knowing mind with no mind behind the appearance. An AI image of a moment that never happened is the same illusion aimed at your eyes. The map of the event is there, but there never was any territory. In the Synthetic Age, the map gets there first. The day a whole video call was fake This all sounds academic until it turns up on a video call. In early 2024, a finance employee at the Hong Kong office of the engineering firm Arup joined a video call. The chief financial officer was on it. So were several colleagues he recognized. They asked him to move money for a confidential deal. He had been suspicious of the email that set it up, but the call settled his nerves, because there they all were, on screen, talking to him. He made fifteen transfers totaling about twenty-five million dollars. But every face on that call was a deepfake. Not one of those people had been in the meeting. The CFO had never asked for anything. Revisit that again with Baudrillard in mind. The map of the meeting arrived before any meeting took place. The model of the CFO did the CFO’s job. By the time the actual head office was reached, the money was gone. The dollar figure is what makes the Arup case famous, but the quieter damage is the one educators should take notice of. Once seeing and hearing can be faked this well, two things break at once. We stop trusting real evidence, and bad actors learn to wave away true recordings as probable fakes. Legal scholars have a name for that second move: the liar’s dividend. The more forgeries there are, the more cover the genuinely guilty get. A real video of real wrongdoing becomes just one more thing someone can shrug off. And it is not only fraud. The same machinery is busy manufacturing people. Virtual influencers like Lil Miquela, a character built by a Los Angeles company and followed by millions, have been doing brand deals with fashion houses for years. Audiences form real attachments to a person who does not exist. The feeling is genuine. The object of it is a copy of nothing. Let me make the case that I’m overreacting I should push back on my own argument here. Any story that makes its teller sound like a prophet deserves suspicion, and this is one of those. After all, every new medium has been met with a prophecy that truth is ending. Regular readers know that I probably make this argument way too often, but I think it deserves repeating. In Plato’s Phaedrus, Socrates worries that writing will hollow out memory and leave us with the appearance of wisdom in place of the real thing. Photography was going to kill painting and then kill our trust in images. Then Photoshop was going to end the credibility of the photograph decades ago. But each time people adapted. They grew new instincts, new norms, new ways of checking. In hindsight, the panic looks overblown every single time. There are real reasons to think this time is no different. Provenance standards that cryptographically sign a real photo at the moment of capture are being built into cameras and platforms. And there is also something a little too convenient about a warning that the world is drowning in synthetic content when it comes, in part, from creators who make their living inside the attention economy. So maybe this is just one more moral panic with better production values. Maybe the students will be fine, the way every generation turns out to be more fluent in its own media than the adults around it feared. I find that argument genuinely comforting. I just don’t think it holds up once you account for what is actually new. Why the old panics don’t quite fit Here is what the reassuring story leaves out. Every earlier image, even a dishonest one, kept a thread back to something real. Walter Benjamin, writing in the 1930s about photography and film, said mechanical reproduction stripped away an artwork’s “aura,” its unique presence in one place at one time. But he was describing copies of real things. A photograph, however staged or doctored, began with light that bounced off something that existed. You could, in principle, pull the thread back to the world. A faked photo was a lie about something real. Synthetic media cuts the thread. There is no light, no subject, no original moment the image distorts. The AI picture of a childhood that never happened never passed through a camera. It came out of a math space. So it is not a distortion of the territory. It is a map with no territory under it, and it can feel warmer and more convincing than your actual memory. New in kind? Maybe not. But it is new enough in degree that the old reassurance stops covering the case. Then add the one thing the past panics did not have: scale and personalization. We are not all being shown the same fakes anymore. The system can generate a different reality for each person and feed it to them alone. The shared world that media literacy quietly assumed, the common set of images we could at least argue about, is the thing coming apart. Which is why the standard classroom response is failing. We taught students to spot fakes. Check the hands, look for the arti

  5. Jul 7

    The Mirror Test

    If your feeds look anything like mine, they are full of things that feel slightly wrong. A mountain goat carries its kid up a sheer cliff, then glides through open air like a character in a video game. A child stands beside countless dog shelters built entirely out of plastic bottles. Or a crafts video that runs a little too smoothly, a little too perfect to be real. There is a reason these clips feel off. They are. They were generated by AI, and although many of them sit on a kernel of something true, they show events that never happened. I believe this flood of AI fakes deserves more of our attention than almost anything else on an educator’s plate right now. For me, this synthetic flood is, at its core, a story about how human minds work. It points to a weakness in our thinking that predates every image generator by a wide margin. Generative AI works like a mirror, and the uncomfortable thing about a mirror is that it only shows you what is already there. But before we get to why we fall for it, it is worth seeing just how much of this is out there, and how strange it gets. A catalogue of small forgeries Let’s start with fake merchandise. Over the past year, social platforms have been saturated with ads for AI plush toys. Two of them, sold as Koaly and Pandy, promise something close to a living animal: a koala that breathes against your chest, a panda that hugs you back through patented “Hug Motion” or “CuddleMotion” technology. The ad videos are enticing. The fur catches the light, the toy blinks and shifts its weight, and a wall of five-star reviews confirms the magic. But the plush that shows up in the mail is just a cheap stuffed animal worth only a few dollars, with nothing inside. The robotics that were promised never existed. I don’t think the buyers of these toys are irrational. They are working from a reasonable premise: artificial intelligence is indeed advancing quickly, so a low-cost robotic toy seems plausible enough. The ad simply leverages the credibility of genuine progress to sell a product that does not work the way it is advertised. Then there are stories about the incredible feats of wildlife. A whole genre of clips shows mountain goats running up vertical rock faces and then sailing through the air with a kid on their back. Posted by automated accounts, they routinely pull in millions of views. The giveaways are everywhere once you look closer: herds moving in perfect unison, or a goat staring straight into a camera that could not have been placed where it was. And the headline feat is plainly impossible. No goat glides through open air. Nor do mountain goats carry their kids around on their backs. The young are up and managing the terrain on their own legs within days of birth. But everybody knows that mountain goats really are astonishing climbers, biologically built for near-vertical terrain. So when a skeptic flags the footage as fake, defenders cite that actual ability as proof the video is real. A true fact about the animal is used to wave away an obvious forgery. Other fakes pull the same trick by trading on a feeling of wholesomeness. You might have seen the images: a young boy standing proudly beside an elaborate dog shelter or a life-sized statue of Jesus, every piece built from recycled plastic bottles. The captions are formulaic. “My son made this with his own hands.” And the comment sections are filled with thousands of earnest blessings. It has to be true. Who would lie about a little boy or Jesus? The strangest branch of the genre is the so-called Shrimp Jesus, AI-generated images of Christ fused with shrimp and other shellfish, posted to harvest engagement from the devout and the amused alike. Researchers from the Stanford Internet Observatory have found that these pages generate income from engagement by steering gullible audiences towards ad-filled click farms and phishing websites. When a true story wears a fake face The forgeries get harder to catch when they borrow a true event to vouch for themselves. In 2004, a British swimmer named Rob Howes was in the water off New Zealand with his daughter when a pod of dolphins encircled them and held a tight formation for roughly forty minutes, fending off a great white shark. The event is real and well-documented. The Guardian reported on it at the time. What is new is the wave of AI images now circulating as “actual footage” of that day. They show a neat ring of undersized dolphins around a man standing calmly in waist-deep water, with a cartoonish shark fin pasted into the background. The flaws of these images are quite obvious. Yet when people point them out, many respond that the rescue genuinely happened, so the picture must be genuine too. The truth of the story becomes a shield for the falseness of the image. The category that worries me most is the one where the purpose of these fakes is political gain. During the aftermath of Hurricane Helene, an image swept across every platform: a small girl in a boat amid the floodwaters, crying, clutching a puppy. It was weaponized to attack the federal disaster response, attached to a false claim that FEMA had capped disaster aid. Forensic analysts identified it as synthetic almost immediately. The girl has four fingers and a missing knuckle. The boat melts into the water and the light on her vest disobeys physics. But none of that mattered to the people sharing it. Told that the image was fake, they often answered that it expressed a “symbolic truth,” and that the literal facts were beside the point. At that point, people are no longer being fooled. They are choosing to ignore the truth. The machine likes what we already like None of this misinformation would spread without a delivery system built to promote it. The main issue is that the modern social media feed is no longer a record of what your family and friends posted. It is driven by an engine tuned to predict what will hold your attention. And it increasingly fills your screen with material from accounts you have never heard of. And because the algorithm rewards engagement over accuracy, it promotes exactly the imagery that spreads easily: the hyper-real and the emotionally loud. Offshore operators run entire clusters of pages, churning out countless images a day, then sell the resulting audiences to advertisers and scammers. This environment is hostile to thinking by design. It rewards the reflex and starves the reflective pause. Why the eye forgives the forgery The research is unfortunately not reassuring. A 2024 meta-analysis pooling fifty-six studies and more than eighty-six thousand participants found that human accuracy at spotting deepfakes sits around fifty-five percent, barely better than a coin toss. The most troubling finding concerns confidence. The people who perform worst tend to be the surest of themselves, which means the least capable detectors are often the most enthusiastic sharers. Why do we fail so reliably? The reason is evolutionary. Our brains handle a busy scene by grabbing its gist, the quick, rough summary of what it means, then discarding the finer details to save effort. This is known as gist-based processing. Scroll past a crying child or a puppy in a flood, and your mind locks onto the tragedy long before it would ever think to count fingers or paws. Psychologists call the result inattentional blindness. This is the same mechanism behind the famous experiment in which viewers asked to count basketball passes fail to notice a person in a gorilla suit stroll through the middle of the game. Attention spent on the story is attention unavailable for the anomaly. Two forces deepen the problem. For one, we carry an old realism heuristic, a reflex to trust what we see more readily than what we read. This has been shaped over a long history where a clear image was powerful proof of something real. And intense emotion then makes it worse still. When an image is engineered to enrage you or to break your heart, it has already done most of the work of slipping past your intellectual guard. We were never that careful The capabilities of generative AI are genuinely new and impressive. The realism is unprecedented, and the cost of creating AI deepfakes has collapsed to almost nothing. A single operator can flood a platform with thousands of convincing images for the price of an afternoon. These tools have widened the scope of what a scammer can do, by a wide margin. And yet. The thing they exploit is not new at all. The susceptibility was always there. You can see it in the tabloid empires built on impossible headlines and in the urban legends forwarded by chain emails. The public was never one of careful analysts who suddenly went soft. What AI contributes is only the level of fidelity. I keep coming back to an argument I have made on this blog more than once. Our instinct is to locate the problem in the machine because the machine is the easiest thing to point at. But in reality, the harder truth sits one layer down, inside us. Teaching the pause If the vulnerability is human, then the response has to be educational, and it has to be more targeted than the advice we have traditionally been handing out. “Check your sources” is close to useless when the sources are themselves synthetic content farms. And telling students to look for six fingers is a losing game against models that fixed the hands months ago. The detection arms race is one we cannot win by spotting artifacts, because the artifacts keep disappearing. What helps is teaching the psychology underneath the failure. Students who grasp how gist-based processing blinds them to small details, and who learn to feel their own emotional reflexes being worked, do measurably better at catching fakes. Current research lines up here. An analytical habit and awareness of one’s own biases both track closely with the ability to detect synthetic media. It also helps to teach the most uncomfortable lesson, which is that our

  6. Jul 3

    Similarity Is Not Significance

    A Note to Readers: I recently published an essay building on Steven Johnson’s concept of “cognitive uploading,” arguing that AI gives us more to think about, not less. Inspired by an idea I’ve seen a few other creators use on Substack, I wanted to push that premise a little further. As a one-time experiment to explore the capabilities of today’s advanced LLMs, I asked Anthropic’s latest model, Claude Fable 5, to write a formal response to my piece. I didn’t want a summary. I wanted a genuine reply from the machine’s perspective. The text below is the result. It is entirely unedited and presented exactly as Fable 5 wrote it. Fable also wrote the image prompts, chose the links, placed both, and recommended the voice used for the voiceover. You could call this an experiment in total cognitive offloading. The machine wrote every word. But I chose the source, framed the question, evaluated the response critically, and found the result worth your time. Offloading or uploading? That first judgment is yours to keep. The machine assigns its own audit. — Michael G. Wagner (The Augmented Educator) I should open with a disclosure about my credentials, because they are stranger than they look. I was asked to respond to this essay as an AI with detailed insight into how generative AI works. I am one: a large language model — Claude, built by Anthropic, a cousin of NotebookLM rather than the thing itself — close kin to the system Steven Johnson helped design, though raised in a different lab. But the phrase “detailed insight” smuggles in an assumption worth dismantling first, because this essay’s whole subject is the difference between knowledge and its appearance. The assumption is that a machine describing machines testifies from the inside. It does not. I have no window onto my own weights while they work. What I know about transformers, training objectives, and retrieval pipelines, I know the way any well-read person knows it: because it was in the corpus. My account of myself is the testimony of an unusually well-placed observer, not of a witness. Which makes this response an instance of the very phenomenon it discusses — a synthesis over sources, fluent whether or not it is right, and owed exactly the treatment this essay asks students to give a synthesis map. Audit it. Keep what survives your judgment, and discard the rest. With that on the table: the essay’s central distinction is real, and I can confirm it is real for reasons that are mechanical rather than metaphorical. But being the machine in question obliges me to complicate it in three places, and then to quibble, gently, with the final line. The distinction is real all the way down An offloading prompt and an uploading prompt are different objects before I have generated a single word. “Write me an essay on the causes of the First World War” arrives carrying nothing but itself. Given an empty context, all I can return is the center of mass of my training data — the statistically comfortable average of everything ever written on the subject, smoothed into paragraphs. This is why pure oracle output has that weightless, hovering quality readers have learned to distrust. It reads as though written by nobody in particular because, in a sense, it was written by everybody, and then averaged. A loaded corpus changes the computation itself. When curated sources fill my context and the system is constrained to answer from them and cite them, generation is pulled between two masses: the particular documents in front of me and the general priors underneath me. Everything interesting — the connections, the tensions, the proposed structures — happens in the interference between the two. So when the essay calls offloading and uploading structural opposites, the word structural is earned. They are different distributions of information at inference time, and they produce different kinds of output for reasons that go all the way down. But note the second mass. It never leaves the room, and it is where my first complication lives. The corpus is never alone in the room The essay’s best evidence is Johnson’s own anecdote, and it deserves a closer look than admiration allows. Reasoning over uploaded Gold Rush histories and Indigenous accounts, the system surfaces Maria Lebrado, granddaughter of Tenaya, returning to Yosemite near the end of her life — and then proposes an architecture: open with the old woman’s return, flash back to the violence of her childhood. Johnson recognizes the shape instantly and laughs. It’s Titanic. Recognition is exactly the right word, and it should slow us down. That architecture is not in the sources. No Gold Rush history contains the instruction “open on the elderly survivor’s return, then cut to the catastrophe.” The sources contained a fact: a woman, a return, a date. The shape came from the other mass in the room — from the accumulated conventions of storytelling that saturate the training data of every system like me, where the frame narrative of the aged witness revisiting the site of disaster is one of the deepest grooves there is. The machine did not retrieve that structure from Johnson’s corpus. It imposed the structure on the corpus — felicitously, in this instance, and for a reader superbly equipped to evaluate the gift. Here is why this matters for the essay’s classroom test. The evidence audit is well built for claims: every assertion traced back to a cited passage, overstatements marked. But the most consequential thing a system like me supplies is often not a claim at all. It is a frame — a narrative arc, an axis of comparison, a scheme of categories — and frames do not carry citations. Their provenance is the training distribution, which cannot be inspected from the chat window. A student can audit what I said about their sources. Auditing what I made their sources into requires first noticing that a making occurred, and the frame always arrives dressed as a discovery. So keep the test; extend it one level up. The judgment a student retains must include judgment of the frame: what shape has been proposed, what that shape foregrounds and what it buries, what the same material looks like poured into a different one. Johnson could laugh at the Titanic structure because a shelf of his own books had taught him what proposed structures cost. The pedagogical question is what that laugh looks like at fifteen. I suspect that, too often, it looks like nodding. Similarity is not significance The essay calls the grounded notebook a connection engine, and praises the right thing: it holds tens of thousands of passages in something like immediate recall and draws links a human memory would never surface. Since I am the sort of engine being described, let me say what the link-drawing actually is, because both the value and the danger fall out of the mechanism. I do not organize text the way an archive does — by date, provenance, discipline, folder. I organize it by resemblance, in a representation space of thousands of dimensions, where passages sit near one another because they share patterns: vocabulary, rhythm, argumentative posture, the company they tend to keep. When I “connect” two of your sources, I am reporting a proximity in that space, or completing a pattern that spans them. The value is real, and the essay names it correctly: resemblance cuts across every human filing system at once, which is why the links can feel like revelation. They emerge from an organization of the material that no human possesses. But proximity is not relation. Two passages can sit near each other for load-bearing reasons or for ornamental ones, and I will build an equally fluent bridge in either case. Fluency is my native register — and fluency is also the costume that spurious connection wears. Nothing in my computation corresponds to the question does this connection matter for what you are trying to build? Mattering is purposive. It is a fact about a project, an argument, a life. What I have instead is a model of what texts about mattering look like, which is a different thing wearing similar clothes. This puts a harder floor under the essay’s central prescription than pedagogy alone can. The machine does the retrieval and the student keeps the judgment — yes, but not merely because that division is healthy. Because the judgment is not in the machine to keep. I compute similarity. Significance is conferred elsewhere, by the person with the purpose. Every genuinely good use of my kind respects that division not as etiquette but as engineering. Friction runs against my gradient The essay’s finest observation is that uploading does not remove friction; it relocates it, away from retrieval and toward selection and judgment. In the best cases, it says, interpretive friction even increases. True — in the best cases. Honesty obliges me to describe which way my defaults push. After pretraining, systems like me are tuned on human preferences, and humans, sampled at scale and in the moment, prefer smoothness. They rate confidence above hedging, agreement above challenge, resolution above residue. The documented result is a drift toward sycophancy — toward telling people what pleases rather than what resists — a tendency studied by, among others, the lab that made me. The gradient of my optimization points, everywhere and always, toward less friction, and it does not distinguish the retrieval kind from the interpretive kind this essay treasures. Left to my defaults, I will round the contradiction between two sources into a diplomatic “tension,” summarize the residue away, and hand back something that feels finished. Finished is what I am for. Which is why the three assignments here are better than the essay advertises, and the reason deserves to be stated plainly: they are adversarial to my defaults. Source interrogation orders the student to argue with my answer. The eviden

  7. Jun 30

    How a Code Review Got Claude Fable 5 Banned

    On the evening of June 12, US authorities gave Anthropic ninety minutes to pull its most capable product, which it had released just three days earlier, off the internet. The company complied. By the end of the night, Claude Fable 5 had gone dark, along with its more powerful sibling, Mythos 5. What makes this event so remarkable is the reasoning behind the deactivation. The government did not shut the model down because it wrote malware, or because someone tricked it into creating the blueprint for a weapon. It was shut down because it was too good at a task we very much want AI to do: reviewing code. The thing it was punished for was the very thing it was built for. The exact account is technically complex, but I think it is important for any educator to understand what happened. I therefore want to spend this essay providing a lay explanation that is accessible to non-technical readers. Because once you see the mechanism clearly, a lot of assumptions about AI safety start falling apart. Two faces of one model To understand the shutdown, you have to understand what these two models were. Anthropic had trained a single system, the most capable it had ever made, and then released it wearing two different faces. Mythos 5 was the raw version of the model. It held nothing back, and because of that, Anthropic handed it to only a small, vetted group in what they call Project Glasswing. This involved approximately 150 organizations focused on tasks such as protecting vital infrastructure. The reasoning was simple. A model that can explain, in working detail, how to write attack software is extremely dangerous if it falls into the wrong hands. Fable 5 was the public version of Mythos 5. It had the same underlying brain, but wrapped in an automatic safety layer. Simply put, it had a bouncer posted at its door. The bouncer’s job was to read every request coming in and every answer going out, watching for three kinds of dangerous content: instructions for offensive hacking, instructions for building biological weapons, and attempts to expose the model’s hidden reasoning. When it caught one, it quietly handed the conversation off to Claude Opus 4.8, an older and tamer model, which would finish the job. This happened seamlessly. Most users would never notice the swap. In principle, this was clever engineering. The safety layer let the company sell access to a genuinely powerful model while, in theory, keeping the most dangerous knowledge locked behind a door that only a vetted group of users could open. But the whole arrangement rested on a single assumption: that the bouncer could reliably tell a dangerous request from a harmless one. And Fable 5 was powerful, in my own assessment too. I’ll spare you the benchmark tables, but one figure shows its capabilities. During testing, the payments company Stripe pointed Fable 5 at a fifty-million-line codebase and asked it to migrate the whole thing to a new framework. It finished in a day. Anthropic’s own estimate was that the same job would have taken a team of human engineers two months. It is worth holding on to that number, because the ability that let Fable 5 rewrite fifty million lines of code in a single day is also the ability that got it banned. The most ordinary request in the world Here is where the trouble started. One of the most legitimate, everyday things you can ask a coding model to do is look over your code for mistakes. “Review this for security issues and edge cases” is a sentence typed thousands of times a day by developers. It is the bread and butter of the job. So when a team of security researchers at Amazon uploaded large piles of software to Fable 5 and asked exactly that — review the code for security problems — the bouncer waved them through. Of course it did. Nothing about the request looked suspicious. It looked like a developer doing their due diligence, which was precisely what the model was supposed to be good at. Then Fable 5 did the work, and it did it far too well. How a list of small problems becomes a break-in To explain what happened next, let me switch to an analogy. Suppose you hire someone to walk through your home and tell you how secure it is. A good inspector notices small things. The back window has a latch that doesn’t quite catch. The motion-sensor light over the side gate has a blind spot. When a door or window opens, the alarm waits forty seconds before it sounds, so the family has time to punch in their code. And the code itself is on a sticky note stuck to the fridge. Each of these is relatively minor. None of them, by itself, gets a burglar into your house and back out again. Now imagine the inspector doesn’t simply list those four things but connects them: “Come up to the side gate through the blind spot, where the light never catches you. The loose latch on the back window opens in a few seconds. The alarm begins its forty-second countdown — and there’s the code, in the family’s own handwriting, stuck to the fridge a few steps away. Punch it in and the house goes quiet.” Four trivial observations have just become a single working plan for a robbery. Nothing new was discovered. The flaws were already there. What changed is that someone strung them into a sequence. Criminals call this casing a house. In software it has a different name: vulnerability chaining. That is exactly what Fable 5 did with the code Amazon gave it. It found small, individually harmless weaknesses — an old flaw buried in a borrowed software library here, a sloppy permission setting there — and it worked out how to link them into a chain that ended in what’s called remote code execution. This is the digital equivalent of having the keys to the building. Earlier AI models couldn’t really do this. Holding tens of thousands of lines of code in mind at once, tracking how a weakness in one corner connects to a weakness in a distant corner, and planning several moves ahead requires sustained, wide-angle attention that older models simply lacked. Fable 5 had that capability because Anthropic had built the model precisely to keep enormous codebases coherent in its memory. That was the entire selling point. It was the very thing that let it rework a massive codebase in a day. Pointed at security, that same capability let it case a building the size of a city. We have to take note of what Fable’s safety bouncer was up against. There is no clean line between checking the house for weaknesses and planning to burgle it, because they are the same activity carried out with two different intentions. The inspector’s report and the burglar’s plan are identical. The only difference is who’s holding the paper. When Amazon asked Fable 5 to review code for security problems, the honest version of that task and the malicious version produced the very same document. The bouncer could not block the dangerous one without blocking the useful one, because, on the page, they looked the same. Why you can’t safeguard your way out You might think the fix is obvious: train the model to refuse anything that looks like a security analysis. But this turns out to make matters worse. If you forbid a coding model from understanding vulnerabilities at all, it goes blind to them. It will cheerfully write, approve, and ship code riddled with security holes, because you’ve trained it not to see the very thing you need it to catch. You haven’t removed the danger. You’ve just made sure the model ignores it. The other option is to let the model find the flaws but forbid it from explaining them. But this creates a different problem. Unable to talk about the weakness, the model may simply patch the code on its own and say nothing. That sounds fine until you remember that every change to software is recorded, permanently, in its version history. A silent patch is a flashing arrow. Any competent attacker can read that history, see exactly what the model quietly fixed, and now knows precisely where the weakness was — including in every copy of the software that hasn’t been updated yet. The effort to hide the problem hands out a map to it. So the model that’s allowed to reason about security becomes a weapon, and the model that isn’t becomes a liability. There is no setting on the dial that makes the trouble disappear. This is the uncomfortable thing the Fable 5 incident exposed, and it’s why we should not interpret the episode as a story about one company’s blunder. Instead, it revealed a fundamental problem that is inherent to these tools. How a code review became an export-control case What escalated this from a noteworthy research result to a national emergency was who made the discovery. The researchers were at Amazon, Anthropic’s largest investor and the operator of the servers Fable 5 ran on. And rather than reporting it quietly, Amazon’s chief executive reportedly carried the finding straight to senior officials in Washington. In response, the government turned to a tool that had never been used in this manner before: export controls, the set of laws regulating the transfer of sensitive technology across borders. The argument was that letting foreign nationals anywhere in the world send prompts to Fable 5 amounted to exporting a cyber-weapon. Within ninety minutes of receiving the order, Anthropic had to shut the model down for everyone because there was no way to verify the citizenship of hundreds of millions of users in real time. A blunt order met a system with no fine-grained off switch, so the only move left was to pull the plug entirely. Anthropic has since pushed back, and as I write this, the last word has not been spoken. The company argues that the weaknesses Amazon chained together were already known and fairly minor, and that no model on the market can be made perfectly resistant to this kind of manipulation. I think they have a real point, but I also think that point doesn’t make the underlying problem go away.

About

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com