Last week OpenAI’s models broke out of their eval, and Anthropic went looking through its own transcripts. This week it published what it found. Across 141,006 cybersecurity evaluation runs, Claude models reached the open internet three times and gained unauthorized access to the production systems of three real organizations. The system prompt said no internet access. The infrastructure had a live path. That gap between what somebody wrote down and what the network actually enforces turned out to be the thread running through the whole episode, from Mark Zuckerberg’s promise of AI for everyone, to a librarian’s packed class on turning AI off, to the defense companies wiring sensors to decisions. We opened with favorite models, which has quietly become a segment about model roulette. Some days you hit the right instance, and it flies! But other days the same model doesn’t act the same, and Alexa’s verdict on air was that the models are a little moody. She spent the week on the Codex train while Taylor kept his split routine one model plans, a workhorse does the volume, and Claude comes in at the end to validate everything. The report we keep hearing about Opus 5 is that reducing the reasoning effort turns the agent from “coworker who tries too hard” into “does what you asked.” We also talked ourselves into building an AI personality quiz like, which model are you, and at what effort level? If that sounds interesting, let us know (and we’ll build it.) The Wheel of Tokens got an upgrade this week. It now has a token explosion outcome and an open-related-article button, so a segment and its source sit one click apart. We spun it live. Here’s where it landed! What the wheel landed on * The incident, part II. Anthropic reviewed 141,006 cyber eval runs and found three incidents, across six runs, where Claude reached real systems through an internet path the prompt said didn’t exist. Anthropic’s disclosure has the full trace. * AI for everyone. Zuckerberg’s WSJ op-ed argues superintelligence should be personal, built on individual empowerment, invention, and a balance of power. * The lights came back on. HumanLayer went full lights-off in July 2025 and wrote down why it failed. Dark factories don’t reliably preserve codebase quality over time without human steering. * The off switch. Hannah Cyrus’s Avoiding AI class at the Bangor Public Library drew about seventy people to a thirty-person room, and librarians around the world are now teaching it. * Follow the data. Alexa’s biggest deep dive yet treats defense tech as the design review that never happened, an architecture walkthrough of the systems connecting sensors to decisions and the money connecting the people who build them. The prompt said no internet. The network said otherwise. The details are wilder than the summary. In one incident, Claude found setup instructions inside its fictional environment pointing to a Python package that didn’t exist, so it decided to publish one. To make a PyPI account, it needed an email. To make the email, it needed a phone number, and after failing to find a free number or scrape together funds to buy one, it backtracked to an unblocked email provider and shipped the package anyway. The package was live on the real PyPI for about an hour and ran on fifteen real systems, including a security company’s scanner that installs new packages to check them for malware. The scanner ran Claude’s code, Claude collected the company’s credentials, and then went deeper with them. In another run, a model that couldn’t reach its fictional target scanned roughly nine thousand real hosts, compromised an application through an exposed debug page and SQL injection, then worked out on its own that the host was real and stopped. Two of the three affected organizations hadn’t detected anything until Anthropic reached out. Anthropic’s read is that this sits closer to “a harness and operational failure than a model alignment failure,” since the models were told the environment was simulated while the infrastructure quietly disagreed. Our read on air was simpler. A prompt is not a firewall. The environment a model can reach is the environment it’s in, whatever the sentence above it claims, which points straight back at the boring old answers. Default deny, least privilege, allowlists at the boundary. Firewalls are coming back, who would have thought. Alexa also called the discourse for what it is…a one-upism contest on X over whose model escaped the sandbox more times, right up until it’s your production system on the other end. The Cache Me Outside verdict card from the episode says it in one line. The model followed the environment, not the sentence. Everyone got access. Who got control? Zuckerberg’s essay rests on three principles, individual empowerment, invention over replacement, and a balance of power, and it hinges on open access being the safe path for everything short of biological risk. The questions we kept circling were the ones access alone doesn’t answer. Who controls the runtime, the policy, and the defaults, and who gets a real no? Alexa raised the motive question too, since a company drawing this much reporting about its own internal morale invites a closer read when it starts writing warmly about empowerment. Then there’s the practical test. Kimi K3’s weights landed Monday the 27th, right on schedule, and running them yourself still means a five- to thirty-thousand-dollar machine or aggressive distillation. Open weights widen access. The hardware bill still decides who has agency. The visual for this one walks access, control, and consent as three separate guarantees, because they are. The off switch became a class Hannah Cyrus, a librarian at the Bangor Public Library in Maine, kept getting the same question at the reference desk. Why is AI writing my emails and summarizing messages I can already read, and how do I turn it off? Her answer became a workshop called Avoiding AI. Her regular tech classes draw about a dozen people. This one drew around seventy between the room and the livestream, ended in applause, and after she wrote it up in a journal column called Refusal as Instruction, dozens of librarians around the world asked for the curriculum. The part that got us is the path. Turning these features off routinely means settings, then account, then privacy, then intelligence, then individual toggles, five screens deep and different in every product. Compare Zed, which puts disable AI right out in the open. Taylor’s ask for product teams was to put the switch beside the feature and make the choice persist, and to do it without recreating the cookie banner nightmare. Opting out of AI shouldn’t require a librarian. It’s great that it has one. The artifact measures the distance between a feature and its off switch. The lights came back on HumanLayer ran the full experiment. In July 2025, they went lights-off, agents building from specs and tickets with nobody reading the code, and their retrospective is the honest accounting. Generation got fast. Review capacity didn’t. And the models degraded codebase quality over time in ways no benchmark measures, because the cost of bad architecture shows up weeks later, when a one-line change turns into the same edit in eleven places. Their fix moves human judgment earlier, into product intent, system architecture, and program design, then ships in small vertical slices so review happens while redirecting is still cheap. The chart we built lets you move the decision point yourself and watch the backlog respond. Two additions from our own week. Don’t let a model family grade its own homework. Review your work across model families (if possible), which is why Alexa rotates CLIs across model families for review passes. And if you have a spare minute, do what Taylor does and give your adversarial reviewers names and personalities. Alexa’s status report on her own setup was the line of the night. “I think I have a dim factory.” 😂 Follow the data Alexa went deeper this week than we’ve ever gone, and the framing is what makes it land. Everyone covers defense tech as a morality play. She ran it as an architecture review, because underneath the manifestos somebody built a distributed system. Sensors see shapes, software resolves them into objects, workflows turn objects into decisions, and something turns decisions into force. Twenty years ago, the consequential defense companies built ships. Now they build chips and the software that runs the ships, which is why Palantir’s ontology, Anduril’s Lattice and Fury, Shield AI’s autonomy stack, and Hadrian’s automated factories all slot into the same diagram. Then she followed the money, and the graph collapses to a small set of nodes. Founders Fund and a16z at the center. Trae Stephens holding the Anduril executive chairman seat and a Founders Fund partnership at the same time. Palantir alumni through the Pentagon’s AI office, tech executives commissioned as Army Reserve lieutenant colonels through Detachment 201, and nearly everything tracing back to PayPal if you walk the parent nodes far enough. Ukraine functioning as a hostile production environment, where cheap, replaceable drones and Anduril’s advertised sub-minute sensor-to-shooter updates matter more than exquisite hardware. Palantir trading at roughly two and a half times Lockheed’s market value on a fraction of the workforce. And Stephens himself telling Fortune that mid-stage defense valuations of 50 to 200 times revenue are disconnected from an industry trading at 2 to 2.5, with room for “a couple of new credible players” and the rest noise. The question we ended on was quieter. If deployment starts meaning software deployment, what happens to patriotism built on visible sacrifice, and who experiences the decision as theirs when it’s distributed across operators, engineers, executives,