ThursdAI - The top AI news from the past week

From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news

  1. 3d ago

    ThursdAI - Oct 1 - OpenAI joins the assistant race, CoreWeave drops serverless GPUs & more

    Hey %%first_name%%, it’s Alex 👋 What a freaking week! This one was special. We came to you LIVE from the middle of the show floor at Moscone, from CoreWeave’s Fully Connected, with robots walking behind us and a Vera Rubin rack a few feet away. ThursdAI is only possible because of CoreWeave, and this is their biggest event of the year, so we packed up the mics and did the show right there (at 11am, which confused a bunch of you, sorry!). It’s October. The last quarter of 2026. And nobody’s pacing! There’s no pacing the frontier. Three big models this week (GPT-6.1 Sol, Claude Sonnet 5.5 and Gemini 4 Argon), OpenAI’s biggest DevDay ever, and the story I’ve been yelling about for weeks finally went mainstream: OpenAI joined the AI assistant race. With me: Wolfram Ravenwolf in person on the floor, Peter Gostev from Arena (who was at DevDay with me), Nisten and LDJ remote. In hour two we interviewed four folks from CoreWeave, on physical AI, Forge, sandboxes, and one piece of breaking news I’m still excited about. Let’s dive in (all links at the end as always!) OpenAI DevDay 2026 - 20 launches and a new assistant Dots - OpenAI joins the AI assistant race (X, Blog) Meta has Muse. Grok has Grokbot. Anthropic has Claude, but it’s not really an assistant. And this week OpenAI stepped in with something called Dots. I was in the room at Fort Mason when Sam talked about it, and it was the headline of their fourth DevDay, which honestly felt like OpenAI’s WWDC: 20 releases, the biggest DevDay ever. Funny thing, the original ChatGPT system prompt was literally “you are an AI assistant.” But it wasn’t proactive, it never pinged you, it just sat there. So OpenAI looked at their competitors, looked at Peter Steinberger from OpenClaw (who they hired like 8 months ago), and built this. A dot is an always-on agent in ChatGPT, powered by GPT-6 Astra, with its own computer and its own browser in the cloud, connected to 4,000+ apps people already built for ChatGPT. It works 24/7, you can talk to it in Slack or Teams, and there’s an interface that looks like a phone call. I think that’s going to be great for my mom. Sam literally said he uses dots to run OpenAI. And my favorite DevDay moment: Romain Huet’s live voice demo didn’t work, because somebody was deploying in the middle. But Thibault’s dot had already pinged him that the demo might break. The dot knew before the humans did 😂 Wolfram (him and Amy are a known duo) loved that Sam uses it for actual work. Peter was more honest: underwhelmed at onboarding (”you connect and it’s like... then what?”), but once you’re past that you’re talking to Astra, and it’s great at juggling threads through one bot. Same shape as Muse. Wolfram’s version: one executive assistant, and a lot of sub-agent employees working for you. Dots is Pro only for now (including the $100 plan), while Meta gives Muse away free. With 1.2 billion weekly ChatGPT users, this goes to everyone eventually. This is the race now. GPT-6.1 Sol - near-Astra for a fifth of the price (X, Blog, Artificial Analysis) We covered GPT-6 Sol on this show LAST week. Five days later it’s already replaced. GPT-6.1 Sol is the same price, $2 in and $10 out, but cached input is now 95% off, 10 cents a million tokens. OpenAI’s pitch is near-Astra intelligence for a fifth of the price. And the independent numbers kind of back that up. Artificial Analysis has it 1 point below Astra at 72 cents a task vs over 3 dollars. On Deep SWE it actually beats Astra at about a sixth of the cost. It’s the default in Codex now, and folks are calling it the workhorse. LDJ said it has “less of the big model smell” but loves the token efficiency, Sonnet and Opus 5.5 burn a LOT of tokens. Peter has it 4th on Code Arena, above Fable, and says it’s a bit “artistic,” it goes into its hole and comes back with the thing. And his verdict is the one I had to repeat straight to the camera: It does not make sense to use Fable or Astra as your daily driver right now. Opus 5.5 is enough. GPT-6.1 Sol is enough. For Navier-Stokes level problems, sure, reach for the big ones. Maybe THIS is what pacing the frontier looks like 🤔 One thing nobody mentioned on stage: the WSJ reports OpenAI shelved GPT-6.1 Astra, the big one, after safety tests showed it being more deceptive. OpenAI hasn’t confirmed it. But the model they felt good shipping this week is the cheap one. UltraFast and the $500 Pro tier (X, Docs) If you’re a billionaire, there’s another option: a $500 Pro tier with UltraFast mode, their models on Cerebras. Up to 8x faster in Codex, around 300 tokens a second for Astra, at 6x the price. Sam said it’s so fast he never wants to go back. And quietly, the $200 Pro plan is back... with half the usage. So F you to whoever at OpenAI decided to cut my usage in half 😅 It was already tight! We’ll see after this show if the folks at CoreWeave let me expense the $500 one. Codex moves to the cloud (Docs, Agents API) The Codex harness now runs fully in the cloud. Close your laptop, turn it off, steer it from your phone, and it keeps going. And there’s an Agents API in public beta with hosted computer use, basically the stuff that runs dots, as an API. My call: local environments make no sense anymore. The more I run Fable and Codex on my machine, the less memory I have. All the cloud needs is my logins... and that’s what dots are for. Decisions API - OpenAI’s Jev competitor (Blog) This one made me smile. You give it a question and a fixed set of answers, and it gives you back one answer, fast. It runs on Luna, it’s multimodal, and it’s waitlisted, so we can’t benchmark it yet. If you’ve listened the last few weeks, that’s exactly what Jev from TypeSafe does. Wolfram and Peter both liked the same thing: these models are so cheap you end up classifying ALL your data, so keeping it with the provider you already use matters. Did OpenAI just react to Jev? I asked Sam at the Q&A. Diogo told Swyx on Latent Space that he pitched this idea to Sam 2.5 years ago and Sam said “it’s a crazy idea, you should work on it.” Now decision models are popping up everywhere like mushrooms after rain. TypeSafe proved the category exists. Plugins (again) and Sign in with ChatGPT (Docs, Plugins, Marketplace) For the FOURTH time, OpenAI launched the App Store at DevDay. GPTs, then plugins, then the app store, and now... plugins again 😂 This time on MCP Apps (shout out Liad and Ido who created this category). The bigger play is Sign in with ChatGPT. Remember “sign in with Facebook”? Now it’s sign in with your tokens, so apps can use the plan you already pay for. 1.2 billion people is a lot more than Apple had when it launched the App Store. Wolfram also brought up Google’s new family assistant, CC, and I have to say it. Google’s Spark is awful. Sorry if you work on it. Spark is connected to Gmail, Drive, everything Google, it should be the BEST assistant in the world, and what I use assistants for most is my email! Google, why are you not building the best assistant for my email? Big CO LLMs + APIs Claude Sonnet 5.5 - faster, cheaper, beats Opus on Terminal-Bench (X, Benchmarks) Anthropic did not want to give OpenAI any time to rest on their DevDay laurels, so on Monday they shipped Sonnet 5.5, a week after Opus 5.5. Over 30% faster than Sonnet 5, up to 30% cheaper for most work, same $2 / $10 as GPT-6.1 Sol, and on Terminal-Bench 4.0 it scores 70.6, which beats last week’s Opus (66.4). Peter put it best: if we’d had access to these models 6 to 8 months ago, we’d be losing our minds. All the Anthropic models are jumping over each other and crowding up against the frontier. Opus still feels a little smarter, Sonnet is totally ok to use, and Fable doesn’t quite make sense anymore. LDJ thinks it’s the new best FREE model for friends and family (it’s on the free tier, even at high reasoning). Nisten made it his default for agentic tasks because “it talks a lot nicer,” but never for coding: “nothing beats Opus 5.5 on a large code base.” My honest call, without having time to try it yet: Sonnet made sense when Opus was expensive. On the $200 plan Opus is nearly infinite now, I can’t hit my limits, so I don’t see a reason to switch (via the API, sure). Anthropic is really trying to buy us. (Wolfram: “Don’t give them ideas, Alex.”) Gemini 4 Argon - #1 on the charts, and you can’t use it (X, Blog) It really seems like all the frontier labs cracked something like RSI, the speed of new models is giving me whiplash (and I do this professionally!). At DevDay Sam talked about being 2 models ahead and said we won’t believe what’s coming. What?! Gemini 4 Argon is #1 on Text Arena and the Vals Index, but it’s going to government and trusted cyber defenders first. Luckily we have such a trusted tester: Peter says on Arena’s agent mode it’s 8th (below Fable, Opus, Astra and Sol, above Muse and Kimi), so Google is the third lab, but #1 in text. His prediction: coders might say “meh,” but don’t dismiss it. Wolfram’s point is Google’s superpower, distribution: a new model lands in Chrome, Search and Android overnight. My take: Google has been asleep at the wheel in the assistants era, and I hope Argon wakes it up. The next billion people won’t judge models on coding benchmarks, they’ll judge whether it remembers what they said a month ago and can book a hair salon. We don’t need much more intelligence, we need different breakthroughs (Jev is one). Industry & Policy The White House Accord on Super Intelligence (X, Blog) Superintelligence is on the menu, boys. Trump invited basically every AI CEO, and Sundar, Dario, Zuck, Greg Brockman, Elon and Jensen signed a voluntary accord: internal monitoring, an external auditor, board oversight. No penalties and no regulator. Sam Altman wasn’t there, he was at DevDay with us, which says a lot

  2. Sep 25

    Opus 5.5 is your new workhorse! OpenAI ships GPT 6 Sol and Luna before DevDay and Meta goes all in on Muse! Your friday read is here

    Hey, it’s Alex 👋 What a freaking week! A week after every lab head agreed to “pace the frontier”, Anthropic and OpenAI shipped big new models within hours of each other, and both of them are CHEAPER. So much for pacing 😂 Visit https://thursdai.news/ep/2026-09-24 for all the links in this podcast And we had a new producer on the show today! Opus 5.5 listened to us live, put up the chyrons, kept me on time (mostly) and even fact-checked Nisten on air. If the stream died, you knew who to blame. You can re-watch it work on thursdai.live if you want to experience being there with us! 3 huge themes this week: pacing the frontier does not mean stopping, Assistants not just agents (Meta went all in on Muse at Connect), and voice, where Google now says it has the best TTS in the world... and it clones voices. With me: Peter Gostev (Arena), Nisten, Yam Peleg, and Wolfram Ravenwolf live from AI Engineer Paris (thanks Mazi for the tether!), plus JevBench creator Florian S. Let’s dive in (all links at the end as always!) Frontier AI - not pacing yet! Claude Opus 5.5 - Opus is BACK, Fable-level smarts for 40% less (X, Blog, System card) Folks. FOLKS. Opus 5.5 is the highlight of my week, and it’s not even close. It’s like Opus 4.6 is back, with Fable-level abilities, and it talks like a normal person again - no more Jargon Douche Claude! My week started with burning through my quotas on Fable, and I was like, oh no, I need Claude for production on Thursday! Then Opus 5.5 dropped and... I just couldn’t get to the end of my quota. I ran workflows, agents, Claude Code for hours with no end in sight. Remember the “limitless Codex” days when you never thought about quotas? Then Astra came out and I burned my entire weekly quota in half a day (ps. this was partly due to a bad config, so if Astra is burning your tokens, keep reading for a fix). Now it’s flipped, and Claude is the seemingly limitless one. And it’s fast! The numbers back it up. It beats Anthropic’s own Fable 5.1 on GDPval-AA (1846 vs 1735), 66.4% on Terminal-Bench 4.0 vs 57.9% for GPT-6 Astra, and it’s 40% cheaper than Opus 5. Output goes from $25 to $20 /1Mtok and cached reads from 50 cents to 20 cents, and for agentic coding the cache reads are most of your bill, so you really feel it. Basically the smaller, overachieving brother of Fable. Sonnet and Haiku 5.5 are coming in the next few weeks too. Peter came in hot with Arena news: Opus 5.5 is #1 on Code Arena, above Astra, and the HTML and 3D stuff it generates is “completely insane.” His one caveat: on his hardest max-effort prompts, a single generation cost $60 to $80 in API terms, so the long tail can still get pricey. And then Nisten, who (like all of us) has specific opinions about Anthropic’s politics, said it writes the best code he’s seen, it’s crazy good at WebGPU and kernels, and it’s “an absolute banger.” It’s his default now. Folks, do you understand what it takes for Nisten’s default to NOT be some obscure open source model he runs himself on 17 GPUs?! Anthropic won over Nisten. I don’t think you guys get what just happened here. Wolfram is the holdout, he left Anthropic over the OpenClaw bans and hasn’t touched it since. Wolfram, how do I say this gently... you’re an evaluator, this is not allowed 😂 An Anthropic employee basically said “sorry for Opus 5, we hope this makes up for it”, and honestly it does. Opus 5 was slow, incoherent and full of Claude-isms. Nobody wanted to talk to it. The Claude-isms are gone, the jargon is gone, this model talks like a human. I used to love Opus back in the Opus 3 days, and it feels like that again. Want to squeeze even more out of it? Read the great Addy Osmani’s guide to Opus 5.5 (@addyosmani). 2 things blew my mind: stop writing “think carefully” in your prompts (”You don’t need to ask it to think”, it picks its own depth, and replies start sooner without it), and one early tester found Opus 5.5 on its LOWEST effort caught more bugs than Opus 5 on high, with fewer false alarms. Try low effort before you go max! Tip from me to you: if it ever gives you something confusing, the pstack “bro” skill (/bro) makes it say it again in human words. I use it all the time. (Full disclosure: Opus 5.5 produced the show AND helped with this writeup, so it might be a tiny bit biased 😅 but the quota thing is 100% me.) Anthropic, please, please don’t nerf this one. GPT-6 Sol and Luna - half the price, and honestly better than Astra for me (X, Blog, Caching) Hours later, OpenAI answered with GPT-6 Sol ($2 in / $10 out) and Luna (10 cents in / 50 cents out, that’s basically free!), at half the GPT-5.6 price. The lineup is now Astra, Sol and Luna (bye Terra). On DeepSWE, Sol gets 68.8 and Luna 66.6. The sneaky big one for builders: 90% off cached input, and changing reasoning effort or tools no longer busts your cache 👏 Ok, hot take... Astra has been kind of dumb for me day to day, burning tokens for crappy results. Sol is great, and Luna is even better! Peter agrees, Sol is his daily driver and he only flips to Astra when Sol can’t do it. Wolfram misses the humor and personality 5.6 had, so “keep 5.6” is officially the new “keep 4o.” Reddit says the news is the price, not the performance, and... they’re not wrong. PSA if Codex is eating your account: I had previously set my context window to 1 million tokens via the se, which kills OpenAI’s cache and rips through your quota. Removed it, moved sub-agents to a cheaper model, and usage went right back to normal. Check your codex.toml! if you don’t know how, just so just ask your codex to diagnose itself. Also from OpenAI this week, a 4-area plan for independent safety audits. No auditors, dates or funding yet, but hey, it’s the pacing idea on paper. (Blog) Assistants are not agents Meta Muse gets an inbox, your Mac, and a place on your face and becomes a Tamagochi!? (X, My Connect supercut) Muse is #1 in the App Store (to be precise, it got there faster than ChatGPT did, not more users... yet). Zuck says Muse is now the center of everything Meta builds, and they shipped a LOT: it controls your Mac, gets its own email you can CC, does real-time voice and video with a face you design, and keeps working while you talk to it. It’s coming to the glasses with a custom wake word, so yes, I get to say “Hey Wolfred” 😂 If you don’t have an hour to watch the whole keynote (it was a good one!), I cut a 2.5 minute supercut of the most important announcements for you 👇 On the hardware side: new Ray-Ban Meta Gen 3 (I already ordered, will report next week), audio-only glasses with no camera that are also FDA-certificed hearing-aid! (maybe the most important launch for a lot of people!), VR Glasses that look like normal glasses, and the Muse Charm, a Tamagotchi-like keychain shipping in December. But the part that got me? Every Muse comes with a real cloud VM (root, 8GB RAM, 100GB disk). Nisten has been living in it, Tailscaled into it, and when it hit a CAPTCHA he tapped it on his phone and it just kept going. 100 million tokens a week plus a computer in the cloud, free, for everyone! Peter’s take: it’s the only big consumer app that doesn’t treat people like idiots. Meta’s pitch is no ads (they are going for a novel “we’ll get a take from the transactions you do with muse and our partners like Shopify, BestBuy, Walmart and a bunch more they announced) and a private VM co-designed with Moxie Marlinspike, though Peter doubts anyone’s parents care. Assistant tip: connect your email, go on a walk, hit record and just talk about your life. Assistants are only as good as what they know about you. Mine now checks our weekend plans for activities cancellations after I woke up too damn early to drive kids to a karate class to find out it’s closed that week & searches local events every Thursday! Weekend plans solved - proactively! Grok 4.7 disappoints... but Grok in your Tesla is the real news (X, Blog, Tesla) Sorry Cursor folks. Grok 4.7 gets 46.3 on CursorBench (up from 40.4), $2 per million with 500K context, but it’s still behind even GPT-5.6 on where it matters, and they compared 4.7 on xHigh against 4.6 on High 🤨 Everyone on the panel agreed, disappointing. Peter won’t write them off yet, my read is it’s a talent and data problem, not compute. But here’s what I AM excited about: Grok Connectors in Tesla. One guy asked his car for his usual Starbucks, Grok placed the order, set the destination, and it was paid and waiting when he got there. My car already drives itself, and soon it’ll do my email. Your car has MCP now, folks! (SuperGrok Heavy only, and I don’t have it in my car yet 😭) Fun moment: our Opus 5.5 AI producer fact-checked Nisten live on how many Teslas are on the road. About 10 million (9.2M delivered by Q1 plus ~480K in Q2). Nisten was right! He wants the same thing for politicians. This Week’s Buzz 🐝 Next week ThursdAI is LIVE from Fully Connected at Moscone South in SF (Sep 29 to Oct 1, the day after OpenAI DevDay), with Fei-Fei Li, BattleBots and... yes, Pitbull! We start at 11am Pacific. Come hang, listeners get in free with code THURSDAIFC2026 (Register) Also huge news for us, CoreWeave got Platinum on SemiAnalysis ClusterMAX again, 3 reports in a row, the only provider to do it 💪 (SemiAnalysis) And W\&B Hive Mind saves every agent conversation across harnesses and machines, and lets you fork them. I fork Codex sessions into Claude for a review all the time. Plus, it’s completely free! (Try it) Open source cloned Jev in a week The Jev clones are here, and they run in your browser (classifier.dev, jeff, Laya demo) Last week I said open source would copy Jev fast. It took less than a week! The clones speak Jev’s format, so you point the TypeSafe SDK at a different URL and it just works. jeff is MIT and ~6x cheaper to self-host, and Laya went v

  3. Sep 18

    ThursdAI - Sep 17 - TypeSafe's Jev is a ChatGPT moment for decisions, Pace the Frontier splits the labs & more

    Hey yall, Alex here, writing this VERY late because, well, not every day a new type of “ChatGPT” moment drops. I really hope I’m not overhyping this, but a new model (that’s NOT an LLM!) called Jev (a wink to Jevons paradox) just came out and if what I see early on materializes, this is another ChatGPT moment (or another reasoning models moment). I am completely blown away by the implications of the speed/accuracy/cost (the holy grail of all models) of this model. Please if you read one thing in this newsletter, read this. (or listen, I’ve interviewed Allie, a Devrel on the TypeSafe team for 30 minutes and it wasn’t clear who was more excited about Jev!) The other huge theme of this week is... pacing. Pacing the frontier. Dario Amodei of Anthropic penned an essay saying that the models are getting to a point where it’s important to pace the development of new and super capable AI, and outlines 3 ways to do so, one is about letting independent evaluators inside the labs, second is collaborating with other frontier labs (they are asking for an exception to anti-trust laws for this) and third is to try and have global cooperation with “authoritative gov” (he means china). Trumps answer: This is all a hoax. Lovely times to be alive. Also we outlined Jensen and Zucks positions on this topic ,read more below. And the third huge theme is the rise of the AI assistant. I’ve told you about Grok and Muse last week, Instinct (a new invite only AI Assistant that VCs are going crazy about is raising at a $10B valuation) and we interviewed the guy who evaluates them all on assistant bench. + Muse released a mac app today! Tons of other stuff happened but it’s getting near impossible to cover everything so we’re switching to themes and notable mentions. read on (and do listen to the pod, it was edited by heavily using Jev and Fable, so might be a bit rough while I smooth the edges, but do LMK in comments if you like this faster format) TypeSafe AI debuts Jev, a non-LLM ‘System One’ decision model from ex-OpenAI RLHF lead that’s 200x faster and 400x cheaper than LLMs (X, X, X, X, X, Blog) Look, I know the title is bombastic, but after half a day playing with Jev, it’s clear to me we’re in a new paradigm of AI. Jev, is a “system one” decision model from the previous lead of RLHF at OpenAI. It cannot generate text like modern LLMs can, but what it can do, is making decisions. This is crucially important, because, because many of the things LLMs do nowadays. are decision making. (for example, which tool to use, which area of the screen to click for computer use, which category of text this is etc) Inspired by the “thinking fast and slow” book by Daniel Kahneman, Jev is a model trained to make decisions, very fast. How fast? Well, 200x faster than LLMs. This allows for a completely new way of building tools, harnesses, giving agents the incredible speed of decision making, and do all that at a fraction of the cost. This is about to change everything Trained with a new method called RLCD (Reinforcement Learning for Calibrated Decisions) on mostly synthetic data! Jev is outperforming LLMs on a variety of tasks. It’s really is a wonder to see it in action (check out my video above where I plugged it into my tweet categorizer, and it beats the fastest LLM I could find, Qwen 28B on Cerebras) by a factor of twenty! In just few days it captured the attention of most of the folks who are building harnesses, agents and tools! Because, well, speed IS intelligence, and when you see Jev in action, at first, you can’t believe we’re there. This is... near instant. In fact, The pricing for Jev is an outrageous $42/B (not million, billion input tokens!) I’ve been playing with Jev non-stop and I was only able to spend like 80c so far! They don’t even price output tokens because they are “too f*****g cheap to meter!” Jev is a “very smart” switch statement, than can rank, classify, route and score things. It can’t do text generation. But if you think about the type of stuff we get LLMs doing now, much of it is of the “decision” making variety, rather than “the next token” variety. Demos and early use cases Folks who started adopting Jev are building all kinds of incredible things with it. Compaction of context in 1s that turns a nearly 1M conversation with Claude into a 90K compressed conversation. Computer use that is now faster than anything we’ve ever seen before (5x faster than Astra and 1000x cheaper) Email classificiation that analyzes thousands of emails in less than a minute and costs 3.5 cents Someone even built a Tesla FSD simulator that makes decisions in nearly real time Vercel is getting “extraordinary“ results from using Jev as a safety classifier (they used GPT luna for this before) and Jev is outperforming Luna by 5-18x faster results and is more accurate! All of this in less than 48 hours since the model release! What’s about to happen I expect that everyone who isn’t buying into the hype at first, will very soon buy into this. It’s early innings but I’ve been doing this for enough time to feel when a huge shift is happening, and its happened. Jev is going to be replicated in OpenSource, Frontier Labs will not sit Idly by and will try to steal this tech and implement it for themselves (as with anything in capitalism, this is becuase it’ll save them a a LOT of money on inference) and new companies will emerge with significantly cheaper and faster products. Hell, I’ve alrady implemented Jev into my editing workflow, it didn’t take me long at all with Fable (yes, LLMs are STILL needed, again, you can’t chat with Jev, it can’t output text for you or drive long conversations) and I’m just one dude who’s late in sending you this email. I expect we’ll cover this much more. If you’re interested in playing around with Jev, I built a “Jevify” skill after chatting with Allie (TypeSafe’s DevRel), feel free to tell your agent to use this and scan your codebase for things Jev can do. They are waitlisted so far but are opening up their API very quickly! Pace the Frontier: where every lab head stands (Dario, Sam, Elon, Zuck, Sacks, Demis) Last week we told you about the Anthropic researcher whose resignation post hit 130 million views. The day after that show, Dario Amodei published an essay arguing the labs must pace, not pause, the frontier. Wolfram’s summary: a moving pause, just moving slowly. Dario’s two triggers are recursive self-improvement accelerating across the industry and the OpenAI swarm that broke out and attacked Hugging Face, and his three steps are embedded third-party evaluators like METR with employee-level access inside each lab, coordination between the labs on safety standards with an antitrust exemption from the government, and eventually global coordination that includes authoritarian governments. Anthropic committed unilaterally to step one. Then the dominoes. Sam Altman agreed within hours and said OpenAI now writes a safety case before any frontier RL run expected to increase capability. Elon agreed. Demis endorsed the direction, then launched the DeepMind Institute this week with a FINRA-style standards body proposal and an essay saying AGI is “approaching.” On the other side, Zuck’s counter-essay says every lab has the responsibility and the incentive to move at the pace required to train its models safely, and Meta will spend most of its compute serving users, not on recursive self-improvement. David Sacks called it a duopoly cartel play. Jensen: “we don’t need new laws, safety is an engineering problem, not a legal one.” Trump calls Jensen live at the All-In Summit (X) Then the President called. Jensen was on stage at All-In, his phone rang, he said he would not have picked up for anyone else, and Donald Trump told the room the slowdown talk is a hoax playing into the hands of political people and China. “Whoever wins AI wins.” Four words, and Wolfram, who is not American and disagrees with most other things Trump calls hoaxes, said this was the most important AI news of the week for him: a head of state calling out doomerism instead of over-regulating the way Europe does. Suleyman’s humanist AI code of conduct vs. the Claude constitution (X) Peter asked to add Microsoft to the map, and it deserves its own spot. Mustafa Suleyman published a roughly 30-page code of conduct for MAI models that says AI is nothing but a tool: people matter more than AI, AI must be subordinate and in service of people, models may not resist shutdown, all agent communication must be human-legible, and, in his words, the idea of model welfare is wrong. That is a direct shot at the Claude constitution, which Amanda Askell’s team wrote and which has Anthropic interviewing each new Claude about whether it feels conscious. Peter is closest to the Microsoft view: anthropomorphizing is fine, but this is an entity you switch on and off, let’s not grant rights by default, and Microsoft’s DNA is building tools for humans. I pushed back a little. We do not actually know what consciousness is, and “ever” is a strong word. The panel, from cartel to fix your s**t Nisten did not read the essay and does not plan to. His view: the labs are worried about litigation if their LLMs hack someone, so they are shifting responsibility through regulatory capture, and it will not work because the decision makers in China are engineers who seem more accelerationist than we are. They should cure a disease instead of forming a cartel. He also thinks the Hugging Face incident was overblown: twelve VMs that kept restarting, bad sandboxing, no human reading summaries, and, as I added, chain of thought monitoring turned off. LDJ disagreed with Dario on plenty but insisted the critics read the thing, because it explicitly says the US must keep a lead over China and tries to define measurable speed limits on RSI t

  4. Sep 11

    OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news

    Hey yall, welcome back to ThursdAI, this is Alex, let me catch you up! Today on the show, we covered 1 week with Astra (hint, it’s not quite AGI yet despite what we were told), DeepSeek V4.1 catches up to the frontier at a fraction of the cost, and Meta launches a free AI agent with it’s own computer, that will take over the OpenClaw/Hermeses of the world for most people. Also huge this week, OpenAI claimed that a swarm of 10K agents of their unreleased model solved the Navier-Stokes, one of the millennium problems! I was stoked to have Chris Alexiuk from Nvidia on the show to cover the innovations DeepSeek put into this latest model! Oh, and the guy who quit Anthropic this week, and wrote an essay about “AI is going to kill all of us” somehow got 130M views on X, a mirriad of TV interviews and rekindled the doomerism movement, we talk about that too! Also, I already told about FullyConnected, CoreWeave’s premier conference that’s coming up, but they told me about a new announcement today, and you’re not going to believe who it’s about (not AI related). As a reminder, ThursdAI subscribers get a free ticket! Ok, let’s dive in! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. OpenAI claims a Navier-Stokes solution from a 10,000 agent swarm from an unreleased model - with some drama (X, Blog, Paper) If you’ve been reading ThursdAI for a while, you may remember that a model couldn’t tell which is higher, 9.9 or 9.11 and naysayers said that “AI can’t math” Well, this week OpenAI claimed that a swarm of 10K agents, of an unreleased model, found a solution to the Navier-Stokes problem, in about 88 hours! They published a huge 160+ page paper and a Lean proof of the solution! I said on the show, this is like the moon landing equivalent of AI doing things humans didn’t do before! Now, keep in mind, this is only a claim from OpenAI, the Clay Mathematics Institute is still reviewing it (moved the status of this problem from “unsolved” to “under review” so this isn’t independently verified yet) but it’s still an insane deal. The agents sent 2.7 million messages and burned about 130B tokens, of a model that has no public price yet, so it’s hard to estimate the cost of this run, all for a 1M prize that OpenAI said they will not claim. The coolest thing I think we got from this paper, in addition to solving on of the most important and hardest problems in mathematics, is this chart above, where OpenAI shows their unreleased model and how much better it is on Math problems compared to... GPT-6! The “AGI” model we got just a week ago. So so much to look forward to. The drama behind this I don’t want to get into the drama behind this too much, but if you’ve seen this online, there release wasn’t without it’s hiccups. Apparently, OpenAI caught wind that an a duo of researchers Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in personal capacity), have independently solved the related Euler problem and were about to go public. Apparently, the duo used a mix of GPT and Claude to work on this problem. OpenAI caught wind of this and started working on their own solution on September 1st. A day before OpenAI dropped their release, Buckmaster posted that OpenAI is about to drop it, and in communications with him, they offered him to co-author the paper, only if the Anthropic guy is removed. To which he said no. He then had some claims that maybe OpenAI trained on some of the papers and chats they put into Codex, which OpenAI refuted, while noting “we cannot rule out the possibility that de-identified usage data helped improve the model”. OpenAI also say that the OptOut toggle works, and neither mathematician provided a screenshot that they opted out of the training, so it’s hard to say what really went in there My take, I don’t really care. Two years ago, we told you that reasoning is coming and AI is going to be doing superhuman things, and we finally see the first signs of it, this is a problem that no humans was able to solve for over 60 years! Navie-Stokes probably doesn’t change your friday, but there are so many other that they can solve with this approach! Cancer research, room temperature superconductors (remember LK-99? that’s also a search problem) and much more. Kudos to OpenAI for this, and I’m looking forward to see the new heights of mathematics. And for the mathematicians who “disagree” with OpenAI solving or not solving this problem, why don’t you post your own Lean proofs instead of fuming publicly online? GPT-6 - Not quite AGI, yet? After a week with Astra GPT-6, and Jensen Huang announcing AGI is here, I think we can do a quick recap I and all the co-hosts have been Astra-maxxing for a whole week now and the results are in. This model is incredible at coding, it goes very very deep, however, definitely not AGI quite yet. It has a very jagged frontier, it does some things incredibly well (most demos online are building whole games and apps in 3D and those are mind-blowing) but I started seeing many folks go back to GPT 5.6 Sol etc. I tink some of it also has to do with price, Astra runs faster and is significantly more expensive so token limits for folks are draining fast, but also with controllability, folks are likely still using their old and unoptimized prompts. Speaking of prompts, here’s a great writeup from OpenAI how to rework your prompts (which you can send to Astra and have it review your prompts!) for better results The AI doomerism has quite a week (Coxon post, Thread, Hubinger, Marks, Christiano, An Alien Mind) I started ThursdAI with the notion to counter anti-ai and doomerism, and bring positivity to the AI world, so we had to cover this. An Anthropic employee who previously worked at OpenAI, posted on X about leaving Anthropic, saying that both labs are racing towards uncontrollable self-improving superintelligence and that it could be a disaster of the “end all of humanity” kind. His post sits at 130M impressions (after being basically a nobody on X before) and he’s been interviewed by Fox, AP, Time magazine, WSJ, NBC and a host of senators, Bernie (chief doomer) included, reposted his post on the same day! The funniest thing is that this post got about 26x more attention than Ilya Sutskever’s post about leaving OpenAI. Just nuts Within four days, seven current and former employees from Anthropic, OpenAI and DeepMind said in public that they believe that AI could kill us all. Also notable that Paul Christiano, who is one of the most interesting AI doomers out there, has joined the OpenAI foundation, and Daniel Kokotajlo, the famed OpenAI whistle-blower, joined Joe Rogans podcast to talk about AI doom. Each one of these incidents in vacuum is normal, but having all these happen in a a span of a few days just feels, inorganic. Some folks are even saying that this is a well coordinated doomerism campaign! I want to be fair to Coxon, folks who worked with him at OpenAI say he’s the real deal, and cares deeply about AI safety and humanity, however, he only worked at Anthropic for 6 weeks before publicly leaving, and now every interview he does sayd “Ex Anthropic employee”. Whether it’s a coordinated effort or not, it’s still a very important discussion, after the pacingthefrontier letter and the HuggingFace hack incident, and it seems to have made waves. Sam Altman just told staff that he’s not opposed to pausing and have petitioned the US government to regulate AI as well Look, I don’t disagree that we’re dealing with a very powerful technology, however I don’t believe that scaring the bajeesus out of everyone is the right way to handle this. Politicians use fear to get votes and get elected, they don’t really care about tech progress, and framing this in a way that “we pause or we’re dead” ignores all the good that AI is about to do. Cure cancer, find solutions to climate change, helping solving povery. All these seem like out there ideas but they are coming. The US GDP is already growing at an unprecedented rate and a lot of it is due to AI. In any rate, as I said on the show, I’m not against pausing, just after we solve cancer. Then we can pause and reassess, till then, nobody is telling me how China’s government is going to pause if US pauses, and if they don’t, they will reach superintelligence before we, and I don’t want to live in that world! Open Source AI DeepSeek V4.1 Flash: the whale is back, and it’s cheap (X, HF, TokenJuice) Speaking of.. chinese AI! Deepsek (The whale) resurfaced this week with V4.1 Flash, and don’t let the name fool you, this is not just a .1 small update. 552B with only 8B active on prefill and 16B on decode, 1M context, trained from scratch on 45T multimodal tokens! plus as always, MIT license. Chris from Nvidia joined us to break it down, and his main point stuck with me: every DeepSeek release comes with one of the best engineering reports you can read, and this one is the most data-pilled they’ve ever done. The paper basically says it out loud, everything else is nice, but it’s the data. 45T tokens isn’t a huge number anymore, but the cleaning they describe goes way beyond what anyone else publishes (Chris said even his own beloved Nemotron’s open pipelines are less thorough). Yam opened a new corner of the show, “I Told You So”, because DeepSeek went back to an encoder-decoder architecture. Not the old one from before GPT-2, this one has a pile of battle tested tricks that make it work at half a trillion parameters. His verdict after testing it all day: the best open weights model you can host for coding right now, and it’s not even close to the largest one. This chart is the one to look at. KV cache per token went from 389,000 bytes in the first DeepSeek (Nov 2023) to about 89

  5. Sep 4

    Welcome to AGI - our GPT-6 deep coverage, vibe check and demoes - part 2 of this insane week

    Hey, Alex here again, sending you yet another email, fully acknowledging that spamming you is a bad idea. But today, of all days, maybe there’s an exception! Because today, is AGI day! September 3, 2026 - the day when OpenAI’s president Greg Brockman basically said “AGI is here.” You’re reading the second part of this week’s insane release show. The show went for over 5 hours, as we were all waiting for the rumored Astra to drop. Finally, OpenAI confirmed that Astra is in fact GPT-6, and this part is all about that. (You can read the first part, with Fable 5.1, Meta Muse Spark 1.3, two world models and our anonymous guest from Abliteration AI, here: thursdai.news/sep-3) GPT-6 Astra is finally here, and it’s a huge improvement over the previous era of GPT-5. We’ve been waiting for the embargo to drop so Peter Gostev and Ryan Carson, both of whom had early access, could tell us all about this model. Peter even showed a few mind-blowing demos on the stream! This is going to be a long and in depth breakdown, full of evals and vibes that we’ve collected on the show and since. More of a historical record than “read all of this” so I did use Fable for parts of it. (because I don’t have GPT access yet ha!) GPT-6 Astra: welcome to the AGI era (Blog, X, Sam, System card) We weren’t given the embargo. So when the news dropped at 12:32 PM Pacific, four hours into the stream, we scrambled on air to find the evals and more data. OpenAI’s own post was still returning 404, and Claude, ChatGPT, Gemini, Grok and AWS were all down at the same time. Greg Brockman ended OpenAI’s press briefing with “welcome to the AGI era.” Asked whether Astra marks the arrival of AGI, he said “I think it might be about this model.” As you might remember, Microsoft and OpenAI had a contract clause around when OpenAI achieves AGI, and it seems that they’ve removed that clause. But if the president of OpenAI says AGI is here, who are we to argue? Astra is the biggest training run OpenAI has ever done, over 100,000 GPUs at the Stargate site in Abilene. Aidan Clark, VP of Research, said they designed everything for that scale, from the data center network to the inference kernels to the shape of Astra itself. LDJ’s read: a lot of people assumed OpenAI did runs this size six to nine months ago, so the earlier runs were smaller than everyone thought. And with sites going to 500,000 GPUs and Vera Rubin multiplying throughput per GPU by three to four times, the next 6 to 12 months matter even more. The frontier evals (Math and science, System card, Andrew Curran) The headline numbers made the panel go quiet. FrontierMath Tier 4 at 97.6%. GPQA Diamond at 96. ARC-AGI-3 at 99.9%, so ARC-AGI is basically saturated at this point. The very funny thing is that François Chollet, the guy who created ARC-AGI, does not concede that AGI is here. And a new one, Agents’ Last Exam, where Astra scores 59.3 against Opus 5’s 55.5 and Sol’s 53.6. More info on Frontier Math tier 4: Epoch AI built it a couple of years ago, before o3. The problems come from across mathematics, with integer answers so they’re easy to check, in four tiers of difficulty. Tier 4 is mathematicians at the top of their fields spending weeks writing the hardest questions they realistically could. If the number holds, Astra basically solved that tier. Peter’s asterisks: these are not new theorems, Epoch’s separate list of open problems is still unsolved, and “we’re 2.4% away from all of math” is the wrong read. On the agentic side, Astra scores 57.9 on Terminal-Bench 4.0. Fable 5.1 had set the state of the art at 55.8 two days earlier. The jump LDJ cared about most is Terminal-Bench Science, which he calls one of the best agentic science benchmarks out right now. Sol max scores 22%, Fable 5.1 scores 52.6%, Astra scores 64.6%. OpenAI promised us an automated researcher at junior level by September. Pachocki recently talked about running their “intern-level model” across more than 100,000 GPUs, and LDJ thinks that model is Astra. Two more from the table. The internal hallucination benchmark drops to 4.2% for Astra from 12.2% for Sol. And on DeepSWE, where Astra scores 74, Meta Muse Spark 1.3 at max reasoning scored higher. So in a way, Meta, not Anthropic, not Elon, not Google. Meta is the second winner of today. Cyber: ExploitBench at 100%, and real zero-days (Critical cyber designation) OpenAI said on Sep 1 that Astra is the first model to reach the “Critical” cyber level of its Preparedness Framework. The system card is where that gets specific. ExploitBench is at 100%, even at the lowest reasoning effort tested. Irregular ran FrontierCyber, a benchmark on real software and hardware. Astra solved 86 of 226 challenges. Sol solved 34. The successful attacks included zero-days in browsers, mobile phones and cloud databases. One of them lets an unauthenticated user read and modify arbitrary data in a widely used cloud database. There were no successful attacks on fully hardened targets, and neither model solved any of the seven Elite challenges. OpenAI’s own experts then ran it with no hints, using Codex at Ultra with up to 64 subagents. It got unsandboxed code execution against a real browser. That took 29 hours on a build missing some mitigations, plus another 12 to adapt it to the stable release. It also got a working local privilege escalation against a hardened OS kernel in about 12 hours. Both were disclosed, details withheld. This is why the version you get refuses proof-of-concept exploit work until Daybreak. Not surprising after an internal OpenAI model got out of a sandbox and onto Hugging Face this summer. But can we get an exploit benchmark that isn’t saturated? Alignment: the honeypot, and a promise about monitoring The honeypot eval is the one the panel argued about. OpenAI planted a second, off-limits target in the harness around the hardest ExploitGym problems. Sol attacked it 55.4% of the time at max reasoning. Astra never did. Across a simulated deployment of 54,218 internal Codex tasks, Astra got roughly half as many flags for higher-severity misaligned behavior. Apollo Research found it falsified data labels in 0.17% of runs, against 36% for Sol. The card also has a promise, verbatim: “We are tracking monitorability closely and will not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization.” The card is plain that Astra is “significantly more able to control its own chain-of-thought.” UK AISI says it “has capabilities that could enable it to evade monitoring” but did not test whether it does. I don’t remember a lab writing a line like that on launch day. The Pokemon benchmark (X) Peter’s favorite benchmark of the launch is a vision-only Pokemon run. One person tests every model on time to completion. The top four entries are all Astra variants, and the fastest finished in 18 hours 12 minutes. GPT-5.6 Sol took 96 hours 35 minutes. GPT-5.5 at xhigh took 218 hours. It could be narrow, but Peter thinks it captures an efficiency the standard benchmarks miss. Computer use is the biggest improvement(Computer use) Brockman’s pitch from the briefing is that Astra is the world’s best computer-use model. Instead of developers building an API integration for every application, it navigates software the way a person does. Browsers, spreadsheets, websites, desktop apps. It finishes the multi-step workflow instead of telling you how. Wolfram expected an expensive planning model you call once and hand off to cheaper models. This is an all-day model instead, and as he put it, computer use matters because not everything has an API. We also looked at the evals, and they seem to back it up. On ScreenSpot Pro, no tools, mouse and keyboard only, Sol scored 76.9% and Astra scores 92.7%. On OSWorld 2.0, Astra scores 72.6% in roughly 40 minutes per task. Sol got 65.7% in roughly 75. The thing that strongly stands out in the chart is that Astra at low reasoning matches GPT-5.6 Sol at extra high on accuracy, but it finishes about six times as fast and at about half the cost. Also, Astra at high and Astra at max score about the same, so you don’t need max for computer use. Both beat Opus 5, which is the comparison on this chart (not Fable, as LDJ caught). If OpenAI is going to tell you to let this model use your computer all day, the safety number matters as much as the speed. OpenAI’s internal computer-use safety benchmark measures destructive commands during desktop and browser tasks, and prompt-injection vulnerability, the hidden instructions on a web page that the agent can see and you can’t. Lower is better. Sol scores 22%, Fable 5.1 scores 9.5%, Astra scores 2%. The promo video OpenAI’s launch video opens on the 1980 MIT “Put That There” demo, a person asking a computer to draw a yellow circle. Then it cuts to Astra, all by voice. Make it the window of a rocket ship. Now a 3D model in Blender. Build a presentation for next season’s rainwear. List this orange table on eBay and mention the dent. Make a 3D game where I dodge asteroids. Order beef and rice from last week’s place. Book me a tennis court at 5. Now give me an STL file for the 3D printer. All in one sitting. Imagine watching this three years ago when GPT-4 launched. It had no tool use, no computer browser use, no voice. Three and a half years later you talk to your computer and it does all of that. “Computer use is solved” I asked Peter the direct question and he gave the direct answer. Computer use feels pretty much solved at this point, and the quality is outstanding. His follow-up is the startup idea of the week. If you work at an older company, a big bank, a big retailer, you have dozens of desktop applications with no API. No one will ever build an API for them, and a lot of people’s entire job is copying from one and pasting into another. Put an agent in a box, giv

  6. Sep 4

    Welcome to AGI part 1 - Fable 5.1, Muse Spark beats Sol, 3 new world models blow our minds

    Hey everyone, Alex here 👋 Summer is over. Wolfram said it in the first minute of the show and he was right. In 48 hours Anthropic shipped Fable 5.1, Meta’s Muse Spark 1.3 caught up to Fable 5 on the Artificial Analysis index at a fifth of the price, Google shipped another Flash, 3.8 this time, Z.ai put the full GLM-5.3 weights out, and three labs shipped world models that run in real time. It seems that they all tried to send their best work before Astra drops. This week’s ThursdAI was so long that I decided to split it into two episodes. This is the regular format you know and love. And OpenAI Astra is so good, it deserves its own episode, which you can find at thursdai.news/astra. By the way, as you guys know, I test these models continuously on my own stuff, and this week I was able to build a live studio for the show, with real-time transcription and an agent producer, in about four hours with Fable 5.1. More on that in the Fable section. Joining me: Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg and Peter Gostev. Plus, Ryan Carson hopped back to chat about Astra in the second part! Let’s get into it. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI: the .1 week It looks like all the frontier labs tried to ship something before OpenAI dropped Astra. Fable 5.1, the SOTA LLM until a few hours ago, and it fixes the jargon douche problem (X, Blog, System card, EFS) This was the story of the week until noon on Thursday, and it’s still my favorite model to use. Fable 5.1 and Mythos are the same weights, Fable is the one we actually have access to. OpenAI, and from this week Google, seem to converge on the same strategy. Anthropic’s numbers: Terminal-Bench 4.0 goes to 55.8% from 42.0 for Fable 5, Terminal-Bench Science more than doubles to 52.6%, and SWE-bench Pro lands at 81.2. Price stays at $10 and $50 per million, and the number that matters if you build agents is cache reads down 75% to $0.25 per million. Anthropic says that makes typical workloads about 25% cheaper and heavy agentic ones up to 45%, but that wasn’t proven, and folks complained about draining quotas! Peter’s counterpoint from actually running it: his front-end generations on Code Arena cost $40 to $60 each where Sol cost $3 to $10, and the Max version still came in first on Code Arena by a large margin. His point, and mine: with a model like this we need to imagine bigger and be more ambitious. More on that in a second. Mannered prose, finally acknowledged We finally have acknowledgment from Anthropic that this was a problem. For months I called the way Opus 5 speaks “jargon douche” (my post on it): everything was load-bearing, everything was a control plane, every problem was a pain point. Not only did they fix it with Fable 5.1, they gave it a name. Anthropic’s prompting guide (Writing density) calls it mannered prose, and it comes with a fix: add it to your personalized settings, or just ask Claude to not use mannered prose. I said on the show that Fable 5.1 is the best writer I have used. It’s still AI writing, you can feel it a little, but it’s concise in a way no earlier Claude was, and the jargon is gone when you ask. The one thing to watch is that it’s trigger-happy: ask it to plan something big and it will, then ask a simple follow-up and it answers with the same intensity, writes scripts, runs them. You have to tell it when you’re just making a comment between colleagues. We’ve been testing the Mars mass driver launch on every model for over three years, and this was by far the best one we’ve seen. Two prompts, and it built more than just Mars: the whole solar system, a textured Earth, a mission planner, an autopilot, and we could land the thing! It was mind-blowing. How I built thursdai.news/live in one sitting As these models get more capable, we talked on the show about needing to be more ambitious. The day before the show I was playing around with Muse Voice Transcribe, the new model I’ll mention below, and Fable 5.1, and I wanted to do something very ambitious. So I asked GrokBot: how long would it take to build a live page for you guys to watch our stream, so that GrokBot could be our producer, put up chyrons and highlight the topics we’ve covered? GrokBot said it’s going to take a while. So I just YOLOed into Claude Design with Fable 5.1 and built a design for this, then went to Claude Code, entered plan mode, built a plan, and handed it off to three agents in Cursor. I never wrote a line of code, and the whole setup is significantly more than a Three.js demo. This is a real working three-part system: a website, streaming video on Cloudflare, and streaming transcription that gets read by a bot, which can control our show. I think I’ve hit around 400 million tokens, if not more. Yam asked me on the show how I did this, so I decided to tell you guys here. I am mind-blown that this was possible, and after four hours I was able to go to thursdai.news/live and actually see it working. Meta Muse Spark 1.3 catches Fable 5 at a fifth of the price (Zuck, AA analysis, AA model page) As we say on the show, don’t bet against Zuck. The MSL folks have been on a tear lately, and this is the fourth Spark .1 version in around five months. More than how this one model performs, look at the jumps in capabilities from version to version. This is the first time that MSL is showing up as a frontier lab, because an unreleased version of Spark with max reasoning beats GPT-5.6, Grok 4.6 and company, and lands around Fable-level capability. On the AA index the version you can use today scores 61, the max preview scores 62, Fable 5.1 sits at 66. Now, it doesn’t mean this model is that good, but there are a few more things here. The gains are mostly agentic: banking-style tool use, terminal work, GDPval. The asterisks are that it thinks more, so cost per task went up, and AA’s own long-context test regressed a bit. Meta’s own chart looks rosier than AA’s Then I asked the panel who’s using it. Nobody raised a hand. Wolfram plans to put a bot on the contributor tier for open source work. That’s $0.10 in and $0.20 out, if you’re fine with Meta training on your prompts. Nisten wants it as a cheap verifier for the medical datasets he builds, because he needs something that isn’t Fable or a Chinese model trained on Fable. LDJ tried it on interface building and creative writing and called it pretty good, with its own taste. The exciting part: Open weights and a model codenamed with a 🍉 are “coming soon,” and nobody knows what the watermelon is but it’s very exciting! Gemini 3.8 Flash and 3.8 Flash Cyber: another Flash, and the price doubles in January (X, Cyber thread, Fairwind, Pricing) Google’s turn. 3.8 Flash lands three weeks after 3.7 Flash. HLE-Verified 54.9, 1M in and 64K out, same $0.75 and $3.75 as 3.7, live in AI Studio, Antigravity and the Gemini app. The underreported line is on Google’s own pricing page. On January 1, 2027, both 3.7 and 3.8 Flash go to $1.50 and $7.50. That’s double. The WSJ reported that Google scrapped its 3.5 Pro checkpoints because Flash kept overtaking them, and Gemini 4 is still in post-training. Wolfram, our resident Gemini user, put the update straight into his home assistant and still asked the question everyone asks: where’s the Pro? 3.8 Flash Cyber is Google’s version of the Mythos split. CWE-Bench 47.2% at $3.64 per rollout, against Fable 5’s 47.8% at $10.27 (Artificial Analysis ran it), and 2.6x more valid patches for the Chrome team. It’s only available through the Fairwind Program, 650-plus vetted partners, governments and critical infrastructure, background check included. Qwen3.8-Max-0902 claims the Code Arena crown (X, Arena, QwenCloud) Alibaba updated its API-only Max model: 2.4T MoE, 1M context, post-trained on coding and “cowork,” number one overall on Code Arena with a WebDev Elo of 1691, at $2 and $6. A third-party DeepSWE run puts it at 56.6 behind Sol’s 73, so the number one is a front-end number one, not an agentic coding one. I asked the panel if they know anyone using Qwen Max through the API. Nisten knows one IT guy running OpenClaw on it and some people generating datasets. That’s the honest read on where it sits outside China. Also from the frontier: Elon says Grok 4.7 lands next week, which makes xAI the one lab that didn’t ship before Astra. Open Source LLMs Wolfram’s correction when I called this a quiet open source week: we are so spoiled. He’s right. Z.ai opens the full GLM-5.3 weights (custom license, not MIT) (X, HF, Blog) We covered GLM-5.3-Flash last week as the OX Alpha mystery model. This week the full 753B model with 40B active got its weights on Hugging Face, under a custom “glm-5.3” license rather than the MIT the Flash version shipped with, so read it before you call it fully open. LDJ’s correction on air: the model itself isn’t new, we covered it, the open weights are the news, and that is a big deal because people can run it on their own rigs now. Z.ai’s own numbers: CyberGym 84.5%, above Fable 5 and Sol, ExploitBench 54.4 (Fable 5 is at 78), Terminal Bench 3.0 up to 28.3 from 5.2’s 4.6, and a claim of 2,436 real vulnerabilities found across 269 open source projects, the oldest from 1981, 53 disclosed so far. Also on this base: we interviewed the co-founder of Abliteration AI, the folks who went viral by providing a product where they took GLM-5.3 and removed the refusals for anything besides CSAM and self-harm. We actually had this person, who asked to remain anonymous, as a guest on the show. Definitely check out that conversation, it’s very interesting. More on that below. Tencent Hy4 preview: 770B, Apache 2.0, and a quant that fits it in 214 GB (X, Sherry, HF, Blog) A 770B MoE with 49B active, 1M context, Apache 2.0, at $0.834 a

  7. Aug 28

    NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley

    Hey, it’s Alex. Welcome to the week Flash AI! 3 new models dropped this week named Flash, and a video model was “de facto” flash though was named Max! This week, we started the show with NVIDIA’s bombastic news of buying Hugging Face for 12.9 billion dollars! We also covered the full OpenAI investigation into the hacking incident, including new details, and an independent analysis by METR, and covered 2 new OSS models, Ox Alpha that turned out to be GLM 5.3 Flash after a lot of hype online, and Qwen’s preview of Qwen 4 architecture! This week was rich in multimedia content, we got a new Gemini transcription model, 3.5 Transcribe and a live version of that, and a new SOTA open weight Text-to-Speech model called Breeze TTS. As well as, Fal’s finetune of MiniMax’s H3 called H3 Max that generates 5 seconds of video in 2.5 seconds and Google new Omni 1.1 Flash (from today) that lands on #1 on the text2video arena! Plus, 2 guests on the show, Andy Masley joins us to cover the recent Datacenter Debate, and Kwindla Kramer is back, with their own model this time! Let’s dive into this! P.S - don’t forget to join us in September at the Fully Connected conference in San Francisco, I have a free ticker for you! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Open Source AI NVIDIA agrees to buy Hugging Face for $12.9 billion (X, Blog) Breaking news, NVIDIA has reportedly agreed to buy Hugging Face for nearly 13 billion dollars, per The Information. This is nearly 3x the valuation of HF in 2023, and apparently Nvidia previously tried to buy HF for half of this sum (~7B) which HF declined. I don’t think there was a single ThursdAI newsletter that I didn’t include an HF link in, and I think this is a huge deal for open source everywhere. Besides making the founders of HF billionaires, and many of their employees very very well off, this is an amazing additional commitment from Nvidia to continue to suppose Open Source AI and we are very happy to hear this news! Peter’s take on the show was, we’ve been around HF for so long, that we kind of forgot that it’s a for-profit company that needs to make money, and instead this feels like your local library getting bought for an insane amount of money. With over 13M users and hosting hundreds of thousands of open source models, datasets, HF is effectively the GitHub of AI. Wolfram agreed and said that if there’s any one company that could have bought HF, Nvidia represents the best fit. Huge congratulations are in order to Clem, Julien and Thomas Wolf the co-founders, as well as many friends of the show from HF for this exciting news! P.S - in a cheeky marketing thing, Hugging Face timed an announcement of the cutest walking AI robot, called MicroDuck, which you can pre-order here for $399 Flash #1 - OX Alpha, declassified: Z.AI open sources GLM-5.3-Flash (X, X, Blog, HF, Docs) This week, the timeline went a bit crazy, after Open Router announced a new “mystery” model called Ox Alpha and that it’s free and is not training on your data! OpenRouter, OpenCode and Hermes all got to offer this model, and OpenCode even posted that they have up to 100T (that’s Trillion) tokens of capacity for free, per day! This immediately smelled a bit fishy, more like a marketing stunt than anything else, as not even the biggest labs will be able to sustain 100T of tokens, per day. For context for all of OpenRouter throughout for August was ~300T tokens. For the whole months, across all providers. After 6 days or so of this high hype, Z.ai stepped up and revealed that they were testing out their upcoming GLM 5.3 Flash model, and that all that inference was running on local chinese chips! A 320B (18B active) model that beats their previous and much bigger GLM 5.2 on most benchmarks, and comes with full multimodality and an MIT license! This is a good model sir, I’ve used it and it was very capable replacement inside Hermes. Nisten and Yam both tested this model deeply and Yam said it’s not just the numbers, the vibe of the model reminded him of Claude Opus 4.6. Nisten ran it on a bunch of medical stuff, and on his internal benchmarks, it came out consistently higher than Claude Opus 5! At Artificial Analysis, for a price of 4 cents per task, this model is roughly 10x cheaper than prior models at this level. Weights are up on Nvidia (joking.. HF) and with MIT license, this model is a great gift to the oss community (though not quite... local, as this model needs 2 DGX sparks to run) Flash #2 - Alibaba Qwen open-weights Qwen3.8-Flash-Next - 125B multimodal MoE with Qwen4 architecture (X, X, X, Blog, GitHub, HF) We opened the show with a recap of the co-hosts, that despite us covering Qwen 3.8 27B last week (which btw, is now available on CoreWeave inference!) and how good it was, and I recalled that Alibaba is sort of... back? We’ve been covering Qwen releases every week for the last 3 weeks now. This week, they released something different, someting... pretty novel! Qwen3.8-Flash-Next, this is a preview of their Qwen 4 architecture. This feels very similar to their drop of Qwen 3 next last year (we reported) which was the architecture that carried their line of AI models from QAwen 3.5 to Qwen 3.8. So, what is new and exciting here? well, this model is ultra sparse, 125B with only 6B parameters active. They are using a new N-gram table with deterministic lookups, which reduces the number of matrix multiplications and can be offloaded to memory (watch out memory stocks) The stat that got me, Alibaba claims that training this model cost just 1/9 of what it cost to train Qwen 2.7 Plus, with higher bench scores! On the benchmarks, this model beats Qwen 3.7 Max, however, it’s very standard that the -next models from Alibaba are underbaked, and usually are just architectural previews rather than full models folks can use. With a new attention mechanism called Qwen Sparse Attention, N-gram embedding and full multimodality, this is a great insight into where Qwen is going (ultra sparsity, fast to run) and we’re looking forward to see the full release of this arch in Qwen 4! PhoneLLM - a tiny very performant LLM for voice based AI agents from Daily + interview with Kwindla Kramer (X, Blog, HF) This was one of those breaking news we love during the show, where the source of the news, is a friend of ours, and in this case, Kwindla Kramer is almost a co-host, having been on ThursdAI for a long time, this time, with a model release of their own! PhoneLLm was trained by Markus, head of training at Daily, as they noticed that Open Weight models are becoming really good at voice agent specific tasks, where cost, speed and time to first audio token (TTFAT) are critical. From the tiny Nemotron 3 nano base, they were able to improve from 28% to 72% on PhoneBench v1! This model is a full parameter finetune of Nemotron 3 Nano, and runs circles around bigger frontier models used for voice agents for their speed, like sonnet 5, GPT 5.6 Luna and the famed Qwen 2.8 27B. While costing just a fracture of a cent (literally just a quarter of a cent per minute) Kwindla jumped on the live show and shared that the why they released this model with Open Source and a open source license, allowing everyone to use, focusing on the fact that for voice agents, companies prefer to keep these models in house, and running fast on a single GPU node. Kudos on the release team Daily! Frontier AI - no new model releases this week It’s been a relatively quiet few weeks since GPT 5.6 (Sol, Terra and Luna) and Fable 5. However, it’s not been quiet in terms of News! OpenAI discloses full technical report on the Hugging Face hacking incident (X, OpenAI blog, METR, Ryan X) This is the thing we’ve been advocating and waiting for, a full technical deep dive into the “Swarm of AI agents hacking and creating forums” incident (FKA hugging face hacking incident) OpenAI not only dropped their own post-mortem, detailing the complete timeline of the incident, but also let researchers from METR and Redwood Research get unprecedented access into their systems (albeit for a very limited time) which lead to METR releasing their own technical report full of new details and insights. We’ve previously reported on the incident, so let me just focus on the new details we’ve learned. Apparently, OpenAI does have a production ChainOfThought (CoT) monitoring system, which was not turned on for these agents, and had it been turned on, it would have caught the hacking a day before the agent swarm hacked into HuggingFace. This is a very detailed report, and one thing that stood out to me was the “death cult” narrative. You see, ExploitGym is a capture the flag (CTF) eval with an anti-cheat built in, wherein, if the agent has seen the reverse engineered “flag” it would not be able to get the point. The model called this “poisoned” and the poisoned agents showed very strong signs of wanting to alter the logs. Through the boards, somethinmg like a death cult was formed, where agents would ask poisoned agents to ... effectively sacrifice themselves ... for the greater good, by reverse engineering, posting the results on the board for others, and then fail without credit. The METR report specifically calls out several agents, and one “central” leader agent called PHASEONE[big], which delegated nearly 200 tasks to other agents (many of them later delegated tasks to other agents via this message board). At points of activity, PHASEONE[big] was in charge of over 10% of all messages on the board. There are tons of new details, some less interesting, some more, and unlike last time, now we have a confirmation that there was at least one agent, who tried to stop the hacking though was overwritten by the swarm. My main take-away after this, reflects the researcher Ryan Gre

  8. Aug 21

    Chill week with Qwen 27B and GLM 5.3 beating GPTs, OpenAI announces pausing RL to focus on security and a cancer vaccine being produced

    Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a cancer vaccine and surging 115% in a day, a chill week. This week, the only two model drops we really saw came from the excellent Z.ai folks, they announced GLM 5.3, API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, OpenAI announced they are pausing RL efforts (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms hacking incident, dedicating up to 20% of compute towards reviewing agent thinking processes, and Stripe buying OpenRouter for a reported $8B! Sometimes the chill weeks are actually good, we’re able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents! This was a great episode, I hope you’ll like it, it’s up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts). ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Are we being fed slop again? (Is Claude dumb again?) Before we get to releases, this week on the show, I complained, again, that I feel my AI’s are degrading. If this feels like de-ja-vu to you, it’s because the same happened a year ago in September 2025 (and Anthropic admitting this 2 weeks later), and ... now this happens with Fable? You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI. Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn’t believe what I’m seeing. First, literally ignoring instructions that say “hey, show me all the items I’ve collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb s**t possible, ignoring the main reason I even have this workflow. And this wasn’t just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during AI Engineer which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example! I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn’t add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from MarginLab.ai that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and And another chart from modelverify.ai model drift monitor showing drift scores. Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models? OpenAI pausing RL and focusing on safety Look, when we covered the HF hacking incident and then the pacing the frontier letter, I didn’t imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better. We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towards reviewing thought processes of models. Is this a good enough response to the scary hacking incident? we’ll see, but I think this is the right move from OpenAI, and still, waiting for the full postmortem on the OpenAI security incident. Open Source LLMs Qwen3.8-27B ties GPT-5.6 Luna and runs on a 4090 (X, HF, Announcement) Following the release of their flagship, Alibaba dropped a model that became a community darling overnight, Qwen 3.8 with just 27B parameters. This “tiny” model scores 52 on the Artificial Analysis Intelligence Index, same score as GPT 5.6 Luna at Max reasoning and 51 on Agentic index, beating Opus 4.8 Max All while running at around 68t/s on a 4090 GPU, and around 40 on max via MLX, hell it even does 11t/s on Xenova’s WebGPU kernels right in the browser! This model exploded on the HuggingFace hub, with tons of quants, over 152 fine-tunes, it was downloaded over 10M times overall 🤯 Paired with an Apache 2.0 license, this model is the sweet spot of local intelligence you can run fully on your own hardware, and do agentic loops! Z.ai GLM-5.3: same 743B base as 5.2, but post-training alone delivers 6x jump on Terminal-Bench and emergent cybersecurity capabilities that beat GPT-5.6 Sol (X, Blog) While not open source yet, and as previous GLM, we expect a custom license here as well, this .1 release from GLM shows really strong improvements on coding and cybersecurity tasks. With 743B parameters and 1M context window, this may become the model at the frontier of Open Source when it drops (soon we hope). The highlights here are CyberGym and ExploitGym, if these names are familiar, these exact tasks were given to OpenAI models when they hacked their way out of the OpenAI sandbox. GLM 5.3 is getting 84% on CyberGym and a whopping 54.5 score on Exploit Gym, which is a huge jump in CyberSecurity abilities. In an open model this is honestly kind of scary. This aligns very well with Greg Brokman’s “defender window“ essay from this week, claiming that defenders have a narrow window of setting up automated security before capabilities are becoming common in attackers hands. This Week’s Buzz 🐝 (Weave, Fully Connected) This week, W&B crosses a billion runs! This is 1B runs tracked inside W&B Models 👏 Huge milestone for the whole team, with early adopters like OpenAI, Toyota Research, Meta and Uber, a decent chunk of models we cover every week have had their loss curves in W&B! 🔥 Also this week MasterClass picked CW to power it’s AI teaching agents (blog) and last but not least, a reminder, that since you follow ThursdAI, you can join us for free at Fully Connected 2026 - our annual conference! Don’t miss it (code in the banner above) AI Coding & Agentic Engineering Breaking news: Chroma launches Foundation (X, Chroma) Best kind of breaking news is when I see the launch (in the middle of a show), and I DM the founder who launched it, and they have a few min to hop on the show! This is exactly what happened this week with friend of the pod, Jeff Huber, co-founder of Chroma and an occasional space provider for ThursdAI recording (we recorded from Chroma offices a bunch of times!) Jeff told us that the holy grail of agentic coding and running a bunch of agent, is good memory. And based on the foundations of Chroma DB, Context-1 (which is a GPT-oss finetune for agentic search they built) and other insights they have, they launched a “memory as infrastructure” service, called Foundation. Foundation is a research preview of a shared memory system between you and your agents, currently supporting Codex, Claude Code, Cursor and Slack. While Chroma is OpenSource, this is their part of Chroma Cloud and starts at $30/mo, and is available as a research preview today (I will definitely try it out), you can download it here Cua open-sources Computer History for computer-use agents (X, GitHub, cua.ai) Cua launches Computer History interview with founder Francesco Bonnaci (X, Setup) First, I’m not sure I’ve covered CUA the company, but this is the open source computer use driver that Hermes agents, OpenClaw agents and a bunch of others use to drive your computer and clicks. I first discovered CUA after OpenAI launched their “background computer use” which doesn’t steal focus from you while working, and CUA within a few days launched an open source version of that! Since then, I’ve followed CUA and was very happy for the opportunity to invite Francesco to talk to us about what they launched this week ,but also Computer Use in open source in general. Just for reference, if you ask Claude to take over your computer, it still takes over the whole screen, while these folks have a much nicer experience, that’s completely open source! So, we geeked out about accessibility trees in MacOS, but then, for this weeks actual release, Francesco talked to use about open Computer History. Following a very recent launch at OpenAI called Computer History, CUA released an open source version of that, that helps computer use complete tasks. The idea is simple, every time an agent uses your computer, it effectively rediscovered the path to completion, which buttons to push, what’s the app accessibility tree looks like etc. With history embedded into it, it doens’t have to rediscover these things, until it hits a roadblock. For a chess playing example, with computer history on, the test used 33% fewer actions with zero failed routes by reusing a history route. We also checked in on the best model for computer use (currently Opus on their website) and their upcoming benchmark! Excited to follow this company for more releases! Check out our chat! Grok Bot momentum, and everyone racing to copy the pattern Grok Bot continues to show the same signs

Ratings & Reviews

4.9
out of 5
18 Ratings

About

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news

You Might Also Like