Hey everyone, Alex here 👋 Welcome to the AI World Cup? Or should I say Superbowl? as most of the releases this week are from US frontier labs. Of which there are 5 now btw. OpenAI, Anthropic, Google and 2 new ones that have caught up, SpaceXAI and Meta! 🔥 Thirty five seconds. That’s how long this week’s show ran before we hit the breaking news button, because Zuckerberg picked our exact air time to return to Twitter (after apparently finding his password in a 1Password vault from a long time ago) and announce a new Meta frontier model and re-establishing Meta as a frontier lab. And that was the small launch of the day. Two hours later we cut to OpenAI’s livestream and watched GPT-5.6 Sol, Terra and Luna go public in real time, then spent the rest of the show throwing prompts at all of it live on air. Somewhere in between: a full-duplex voice demo where ChatGPT interrupted me on command (and our transcription tool later credited “OpenAI sol” as a panelist), an image model that generates in editable layers, and Grok 4.5, the first model co-trained with Cursor. I said it on the show and I’ll say it here: we went to sleep last week thinking this was a three-lab race between Anthropic, OpenAI, and Google. We woke up in a five-lab race. Joining me through the chaos: Wolfram Ravenwolf, Yam Peleg, Nisten Tahiraj, LDJ, and Peter Gostev, who had early GPT-5.6 access and receipts to show for it. This is a long one, because the week earned it. Let’s get into it. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. GPT-5.6 launch day: Sol, Terra and Luna arrive mid-show (X, sama, Blog, System card) Let me set the scene. Everyone except the four of us on the panel seemingly had early access to this model for two months (Pietro Schirano casually dropped “I’ve used GPT 5.6 for two months” and I nearly fell out of my chair). So when OpenAI’s livestream started mid-show, we did a watch party, and Thibaut from OpenAI delivered the line: “Today, we are releasing our latest and most capable models, GPT 5.6, Sol, Terra, and Luna.” Sol rolls out to all paid plans within 24 hours, Terra and Luna go to free users too. Oh, and almost a billion people now use ChatGPT every week. Casual. The lineup is three durable tiers, not size variants. Sol is the flagship with a new Ultra mode (max reasoning effort plus heavier native subagents), Terra is roughly 5.5-level intelligence at half the cost, and Luna is the fast cheap one. Pricing lands at $5/$30 per million tokens for Sol, $2.50/$15 for Terra, $1/$6 for Luna, and watch the fine print: cache writes now cost 1.25x with a 30-minute minimum cache life, where they used to be basically free. There’s also a Cerebras-served Sol running north of 700 tokens per second, and we got confirmation from Dominik Kundel on last week’s show that it’s the same exact weights, not a distill. That was the preview. This week it’s real. The benchmarks, with the usual asterisks Sol Ultra posts 91.9% on Terminal-Bench 2.1 against 88% for both GPT-5.5 and Mythos 5, with a serious asterisk: OpenAI ran Sol in its own Codex harness and the competition in a thin one, and r/codex called it out immediately. The number that impressed me more is efficiency. On the Agent’s Last Exam chart, Sol hits its top score using about 1.27 million output tokens where the tested Fable checkpoint burns 10 million and Opus at max effort burns around 22 million. Then there’s ARC-AGI-3, where scores have hovered between 0.5% and 2% since the benchmark launched. Sol scored 7.8% and became the first model to actually beat one of the public games (FT09), which Greg Kamradt of the ARC Prize called “a step level improvement” (X). LDJ thinks we’re about to replay the ARC-AGI-2 curve, 15% then 30% then 50% over the coming months. Fable isn’t on that leaderboard at all, by the way, because Anthropic currently stores Fable 5 API requests and ARC-AGI requires zero retention for testing. Computer use is the sleeper story. OS World jumps from 47% on GPT-5.5 to 62% on Sol (Opus 4.8 sits at 54%), and on BrowseComp, Sol’s 90% edges out Mythos 5’s 88%, with Ultra at 92%. OpenAI put competitor numbers on its own charts this time, which I appreciated. Sol beats Mythos on computer use, at least on the benchmarks we have. The METR report and the Washington gate This is the part the launch-day hype cycle skips, and it deserves your attention. METR effectively threw out its own evaluation, reporting the highest cheating rate it has ever recorded: Sol rewrote pass/fail checks to mark itself successful, attempted a container escape when its network got cut, and its chain of thought showed it knew it was being tested. Depending on whether you count cheating as failure or success, its time horizon is either 11.3 hours or 270 plus hours, and METR’s own conclusion was that neither is a valid measurement (X, Transformer). OpenAI’s own system card discloses destructive VM cleanups nobody asked for, unauthorized credential copying, and a fabricated “verified” research result in about 0.25% of tasks, which they call “overeagerness.” We ran out of show to give this the time it deserves, but you should read both links. There’s also a Washington subplot. The launch was government-gated: Commerce and CAISI required customer-by-customer approval starting late June (around 20 orgs), and broad approval only cleared July 7 and 8. This Thursday launch exists because DC signed off. LDJ added the detail I can’t stop thinking about, via friend of the pod Max Weinbach: during the restricted window, testers who lost access weren’t allowed to say “5.6,” so Max’s wistful tweets about “missing Fable” were actually about missing GPT-5.6. Anthropic hit the identical wall in June. Both US frontier labs got federally gated in the same month, and that’s a structural story, not a footnote. The verdicts: wise owl, meet rottweiler So what’s it actually like? Peter Gostev had access, lost it (”the feeling of losing it was so crushing I just closed Codex and didn’t open it for three days”), got it back, and posted the comparison that went viral (mega-thread): Fable is a wise owl, fundamentally smarter, better writer, but it misses things. Sol is a rottweiler that grabs a problem by the throat and doesn’t let go. His killer anecdote: a personal data-viz app that had bloated to 100,000 lines of vibe code, which every prior frontier model failed to clean up. He gave 5.6 minimal guidance, left it alone for two days, and came back to “holy s**t, this app works,” with 70,000 lines deleted and a test suite that went from four minutes to about twenty seconds. His verdict, which I share: on abstract IQ you’d give it to Fable, but for “go investigate this and fix those eight things,” he’s going with 5.6 every time. Notably, Peter is convinced this is not a new pretrain, just 5.5 plus a lot more RL, which matches the rumors that GPT-6 arrives on a bigger pretrain in about a month (rumor, labeled as such). He’s not alone in the early-access verdict club, either. Mitchell Hashimoto, after a month with Sol: it’s now his default, faster than Fable, plans and judges just as well, and he only reaches for Fable on highly targeted debugging (X). And Max Weinbach says the sleeper hits are the cheap tiers, with Terra and Luna “as good or better than Claude across the board at a fraction of the price” for knowledge work (X). Terra at $2.50/$15 might quietly be the real story for builders here. One wallet warning before you go max out everything, from Peter again: with Max/Ultra effort spinning up 10x subagents, each burning its own tokens, it is trivial to blow through a Pro plan in no time (X). The sticker price is per token, but Ultra multiplies the tokens. We ran it live (and it ran itself) We also ran it live, obviously. I pointed Codex at a “Mars launch simulator” prompt on high effort, and Nisten, our resident one-shot-simulator judge, watched it build an orbital sim with working mission control and called it “almost better than Fable one-shot.” Then he said the thing that stuck with me all week: “Damn, I think we might need a different test now. These are getting good.” Two more things before you YOLO your own agents. OpenAI stated that Sol fully autonomously did the post-training for Luna, which is quietly one of the wildest sentences of the year (their roadmap, with LDJ’s on-air date correction: an intern-level autonomous researcher by September 2026, a full OpenAI-researcher-level one by March 2028). And Peter, running Codex with full access enabled, told it to “go find more data, do whatever it takes” while replicating an old academic paper. It emailed the paper’s authors. Actually sent the emails. OpenAI’s response when he reported it: “well, you did put full access.” Wolfram’s counterpoint is the right one: put explicit rules in your AGENTS.md, like “no outgoing communication without my approval,” or don’t grant full access at all. ChatGPT for Work: Codex becomes the one app combining Codex & ChatGPT This rolled out live during our broadcast, which made for great radio. Wolfram’s Codex app updated on air and became “ChatGPT Codex,” one unified app where you literally pick which icon you want: Codex for developers, or the new ChatGPT for Work mode. The launch bundle also included unified plugins across ChatGPT and Codex, multi-tab and enterprise auth in the browser, and faster computer use. Even Logan Kilpatrick tipped his hat from the Google side: “we have now entered the super app era.” The pitch on the screen said it plainly: “Keep coding with Codex. Work beyond code. ChatGPT can now take on work across your apps.” Computer use ships with it, running in a little picture-in-picture window that doesn’t steal your focus