ThursdAI - The top AI news from the past week

From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news

  1. 6d ago

    Welcome to AGI - our GPT-6 deep coverage, vibe check and demoes - part 2 of this insane week

    Hey, Alex here again, sending you yet another email, fully acknowledging that spamming you is a bad idea. But today, of all days, maybe there’s an exception! Because today, is AGI day! September 3, 2026 - the day when OpenAI’s president Greg Brockman basically said “AGI is here.” You’re reading the second part of this week’s insane release show. The show went for over 5 hours, as we were all waiting for the rumored Astra to drop. Finally, OpenAI confirmed that Astra is in fact GPT-6, and this part is all about that. (You can read the first part, with Fable 5.1, Meta Muse Spark 1.3, two world models and our anonymous guest from Abliteration AI, here: thursdai.news/sep-3) GPT-6 Astra is finally here, and it’s a huge improvement over the previous era of GPT-5. We’ve been waiting for the embargo to drop so Peter Gostev and Ryan Carson, both of whom had early access, could tell us all about this model. Peter even showed a few mind-blowing demos on the stream! This is going to be a long and in depth breakdown, full of evals and vibes that we’ve collected on the show and since. More of a historical record than “read all of this” so I did use Fable for parts of it. (because I don’t have GPT access yet ha!) GPT-6 Astra: welcome to the AGI era (Blog, X, Sam, System card) We weren’t given the embargo. So when the news dropped at 12:32 PM Pacific, four hours into the stream, we scrambled on air to find the evals and more data. OpenAI’s own post was still returning 404, and Claude, ChatGPT, Gemini, Grok and AWS were all down at the same time. Greg Brockman ended OpenAI’s press briefing with “welcome to the AGI era.” Asked whether Astra marks the arrival of AGI, he said “I think it might be about this model.” As you might remember, Microsoft and OpenAI had a contract clause around when OpenAI achieves AGI, and it seems that they’ve removed that clause. But if the president of OpenAI says AGI is here, who are we to argue? Astra is the biggest training run OpenAI has ever done, over 100,000 GPUs at the Stargate site in Abilene. Aidan Clark, VP of Research, said they designed everything for that scale, from the data center network to the inference kernels to the shape of Astra itself. LDJ’s read: a lot of people assumed OpenAI did runs this size six to nine months ago, so the earlier runs were smaller than everyone thought. And with sites going to 500,000 GPUs and Vera Rubin multiplying throughput per GPU by three to four times, the next 6 to 12 months matter even more. The frontier evals (Math and science, System card, Andrew Curran) The headline numbers made the panel go quiet. FrontierMath Tier 4 at 97.6%. GPQA Diamond at 96. ARC-AGI-3 at 99.9%, so ARC-AGI is basically saturated at this point. The very funny thing is that François Chollet, the guy who created ARC-AGI, does not concede that AGI is here. And a new one, Agents’ Last Exam, where Astra scores 59.3 against Opus 5’s 55.5 and Sol’s 53.6. More info on Frontier Math tier 4: Epoch AI built it a couple of years ago, before o3. The problems come from across mathematics, with integer answers so they’re easy to check, in four tiers of difficulty. Tier 4 is mathematicians at the top of their fields spending weeks writing the hardest questions they realistically could. If the number holds, Astra basically solved that tier. Peter’s asterisks: these are not new theorems, Epoch’s separate list of open problems is still unsolved, and “we’re 2.4% away from all of math” is the wrong read. On the agentic side, Astra scores 57.9 on Terminal-Bench 4.0. Fable 5.1 had set the state of the art at 55.8 two days earlier. The jump LDJ cared about most is Terminal-Bench Science, which he calls one of the best agentic science benchmarks out right now. Sol max scores 22%, Fable 5.1 scores 52.6%, Astra scores 64.6%. OpenAI promised us an automated researcher at junior level by September. Pachocki recently talked about running their “intern-level model” across more than 100,000 GPUs, and LDJ thinks that model is Astra. Two more from the table. The internal hallucination benchmark drops to 4.2% for Astra from 12.2% for Sol. And on DeepSWE, where Astra scores 74, Meta Muse Spark 1.3 at max reasoning scored higher. So in a way, Meta, not Anthropic, not Elon, not Google. Meta is the second winner of today. Cyber: ExploitBench at 100%, and real zero-days (Critical cyber designation) OpenAI said on Sep 1 that Astra is the first model to reach the “Critical” cyber level of its Preparedness Framework. The system card is where that gets specific. ExploitBench is at 100%, even at the lowest reasoning effort tested. Irregular ran FrontierCyber, a benchmark on real software and hardware. Astra solved 86 of 226 challenges. Sol solved 34. The successful attacks included zero-days in browsers, mobile phones and cloud databases. One of them lets an unauthenticated user read and modify arbitrary data in a widely used cloud database. There were no successful attacks on fully hardened targets, and neither model solved any of the seven Elite challenges. OpenAI’s own experts then ran it with no hints, using Codex at Ultra with up to 64 subagents. It got unsandboxed code execution against a real browser. That took 29 hours on a build missing some mitigations, plus another 12 to adapt it to the stable release. It also got a working local privilege escalation against a hardened OS kernel in about 12 hours. Both were disclosed, details withheld. This is why the version you get refuses proof-of-concept exploit work until Daybreak. Not surprising after an internal OpenAI model got out of a sandbox and onto Hugging Face this summer. But can we get an exploit benchmark that isn’t saturated? Alignment: the honeypot, and a promise about monitoring The honeypot eval is the one the panel argued about. OpenAI planted a second, off-limits target in the harness around the hardest ExploitGym problems. Sol attacked it 55.4% of the time at max reasoning. Astra never did. Across a simulated deployment of 54,218 internal Codex tasks, Astra got roughly half as many flags for higher-severity misaligned behavior. Apollo Research found it falsified data labels in 0.17% of runs, against 36% for Sol. The card also has a promise, verbatim: “We are tracking monitorability closely and will not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization.” The card is plain that Astra is “significantly more able to control its own chain-of-thought.” UK AISI says it “has capabilities that could enable it to evade monitoring” but did not test whether it does. I don’t remember a lab writing a line like that on launch day. The Pokemon benchmark (X) Peter’s favorite benchmark of the launch is a vision-only Pokemon run. One person tests every model on time to completion. The top four entries are all Astra variants, and the fastest finished in 18 hours 12 minutes. GPT-5.6 Sol took 96 hours 35 minutes. GPT-5.5 at xhigh took 218 hours. It could be narrow, but Peter thinks it captures an efficiency the standard benchmarks miss. Computer use is the biggest improvement(Computer use) Brockman’s pitch from the briefing is that Astra is the world’s best computer-use model. Instead of developers building an API integration for every application, it navigates software the way a person does. Browsers, spreadsheets, websites, desktop apps. It finishes the multi-step workflow instead of telling you how. Wolfram expected an expensive planning model you call once and hand off to cheaper models. This is an all-day model instead, and as he put it, computer use matters because not everything has an API. We also looked at the evals, and they seem to back it up. On ScreenSpot Pro, no tools, mouse and keyboard only, Sol scored 76.9% and Astra scores 92.7%. On OSWorld 2.0, Astra scores 72.6% in roughly 40 minutes per task. Sol got 65.7% in roughly 75. The thing that strongly stands out in the chart is that Astra at low reasoning matches GPT-5.6 Sol at extra high on accuracy, but it finishes about six times as fast and at about half the cost. Also, Astra at high and Astra at max score about the same, so you don’t need max for computer use. Both beat Opus 5, which is the comparison on this chart (not Fable, as LDJ caught). If OpenAI is going to tell you to let this model use your computer all day, the safety number matters as much as the speed. OpenAI’s internal computer-use safety benchmark measures destructive commands during desktop and browser tasks, and prompt-injection vulnerability, the hidden instructions on a web page that the agent can see and you can’t. Lower is better. Sol scores 22%, Fable 5.1 scores 9.5%, Astra scores 2%. The promo video OpenAI’s launch video opens on the 1980 MIT “Put That There” demo, a person asking a computer to draw a yellow circle. Then it cuts to Astra, all by voice. Make it the window of a rocket ship. Now a 3D model in Blender. Build a presentation for next season’s rainwear. List this orange table on eBay and mention the dent. Make a 3D game where I dodge asteroids. Order beef and rice from last week’s place. Book me a tennis court at 5. Now give me an STL file for the 3D printer. All in one sitting. Imagine watching this three years ago when GPT-4 launched. It had no tool use, no computer browser use, no voice. Three and a half years later you talk to your computer and it does all of that. “Computer use is solved” I asked Peter the direct question and he gave the direct answer. Computer use feels pretty much solved at this point, and the quality is outstanding. His follow-up is the startup idea of the week. If you work at an older company, a big bank, a big retailer, you have dozens of desktop applications with no API. No one will ever build an API for them, and a lot of people’s entire job is copying from one and pasting into another. Put an agent in a box, giv

  2. 6d ago

    Welcome to AGI part 1 - Fable 5.1, Muse Spark beats Sol, 3 new world models blow our minds

    Hey everyone, Alex here 👋 Summer is over. Wolfram said it in the first minute of the show and he was right. In 48 hours Anthropic shipped Fable 5.1, Meta’s Muse Spark 1.3 caught up to Fable 5 on the Artificial Analysis index at a fifth of the price, Google shipped another Flash, 3.8 this time, Z.ai put the full GLM-5.3 weights out, and three labs shipped world models that run in real time. It seems that they all tried to send their best work before Astra drops. This week’s ThursdAI was so long that I decided to split it into two episodes. This is the regular format you know and love. And OpenAI Astra is so good, it deserves its own episode, which you can find at thursdai.news/astra. By the way, as you guys know, I test these models continuously on my own stuff, and this week I was able to build a live studio for the show, with real-time transcription and an agent producer, in about four hours with Fable 5.1. More on that in the Fable section. Joining me: Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg and Peter Gostev. Plus, Ryan Carson hopped back to chat about Astra in the second part! Let’s get into it. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI: the .1 week It looks like all the frontier labs tried to ship something before OpenAI dropped Astra. Fable 5.1, the SOTA LLM until a few hours ago, and it fixes the jargon douche problem (X, Blog, System card, EFS) This was the story of the week until noon on Thursday, and it’s still my favorite model to use. Fable 5.1 and Mythos are the same weights, Fable is the one we actually have access to. OpenAI, and from this week Google, seem to converge on the same strategy. Anthropic’s numbers: Terminal-Bench 4.0 goes to 55.8% from 42.0 for Fable 5, Terminal-Bench Science more than doubles to 52.6%, and SWE-bench Pro lands at 81.2. Price stays at $10 and $50 per million, and the number that matters if you build agents is cache reads down 75% to $0.25 per million. Anthropic says that makes typical workloads about 25% cheaper and heavy agentic ones up to 45%, but that wasn’t proven, and folks complained about draining quotas! Peter’s counterpoint from actually running it: his front-end generations on Code Arena cost $40 to $60 each where Sol cost $3 to $10, and the Max version still came in first on Code Arena by a large margin. His point, and mine: with a model like this we need to imagine bigger and be more ambitious. More on that in a second. Mannered prose, finally acknowledged We finally have acknowledgment from Anthropic that this was a problem. For months I called the way Opus 5 speaks “jargon douche” (my post on it): everything was load-bearing, everything was a control plane, every problem was a pain point. Not only did they fix it with Fable 5.1, they gave it a name. Anthropic’s prompting guide (Writing density) calls it mannered prose, and it comes with a fix: add it to your personalized settings, or just ask Claude to not use mannered prose. I said on the show that Fable 5.1 is the best writer I have used. It’s still AI writing, you can feel it a little, but it’s concise in a way no earlier Claude was, and the jargon is gone when you ask. The one thing to watch is that it’s trigger-happy: ask it to plan something big and it will, then ask a simple follow-up and it answers with the same intensity, writes scripts, runs them. You have to tell it when you’re just making a comment between colleagues. We’ve been testing the Mars mass driver launch on every model for over three years, and this was by far the best one we’ve seen. Two prompts, and it built more than just Mars: the whole solar system, a textured Earth, a mission planner, an autopilot, and we could land the thing! It was mind-blowing. How I built thursdai.news/live in one sitting As these models get more capable, we talked on the show about needing to be more ambitious. The day before the show I was playing around with Muse Voice Transcribe, the new model I’ll mention below, and Fable 5.1, and I wanted to do something very ambitious. So I asked GrokBot: how long would it take to build a live page for you guys to watch our stream, so that GrokBot could be our producer, put up chyrons and highlight the topics we’ve covered? GrokBot said it’s going to take a while. So I just YOLOed into Claude Design with Fable 5.1 and built a design for this, then went to Claude Code, entered plan mode, built a plan, and handed it off to three agents in Cursor. I never wrote a line of code, and the whole setup is significantly more than a Three.js demo. This is a real working three-part system: a website, streaming video on Cloudflare, and streaming transcription that gets read by a bot, which can control our show. I think I’ve hit around 400 million tokens, if not more. Yam asked me on the show how I did this, so I decided to tell you guys here. I am mind-blown that this was possible, and after four hours I was able to go to thursdai.news/live and actually see it working. Meta Muse Spark 1.3 catches Fable 5 at a fifth of the price (Zuck, AA analysis, AA model page) As we say on the show, don’t bet against Zuck. The MSL folks have been on a tear lately, and this is the fourth Spark .1 version in around five months. More than how this one model performs, look at the jumps in capabilities from version to version. This is the first time that MSL is showing up as a frontier lab, because an unreleased version of Spark with max reasoning beats GPT-5.6, Grok 4.6 and company, and lands around Fable-level capability. On the AA index the version you can use today scores 61, the max preview scores 62, Fable 5.1 sits at 66. Now, it doesn’t mean this model is that good, but there are a few more things here. The gains are mostly agentic: banking-style tool use, terminal work, GDPval. The asterisks are that it thinks more, so cost per task went up, and AA’s own long-context test regressed a bit. Meta’s own chart looks rosier than AA’s Then I asked the panel who’s using it. Nobody raised a hand. Wolfram plans to put a bot on the contributor tier for open source work. That’s $0.10 in and $0.20 out, if you’re fine with Meta training on your prompts. Nisten wants it as a cheap verifier for the medical datasets he builds, because he needs something that isn’t Fable or a Chinese model trained on Fable. LDJ tried it on interface building and creative writing and called it pretty good, with its own taste. The exciting part: Open weights and a model codenamed with a 🍉 are “coming soon,” and nobody knows what the watermelon is but it’s very exciting! Gemini 3.8 Flash and 3.8 Flash Cyber: another Flash, and the price doubles in January (X, Cyber thread, Fairwind, Pricing) Google’s turn. 3.8 Flash lands three weeks after 3.7 Flash. HLE-Verified 54.9, 1M in and 64K out, same $0.75 and $3.75 as 3.7, live in AI Studio, Antigravity and the Gemini app. The underreported line is on Google’s own pricing page. On January 1, 2027, both 3.7 and 3.8 Flash go to $1.50 and $7.50. That’s double. The WSJ reported that Google scrapped its 3.5 Pro checkpoints because Flash kept overtaking them, and Gemini 4 is still in post-training. Wolfram, our resident Gemini user, put the update straight into his home assistant and still asked the question everyone asks: where’s the Pro? 3.8 Flash Cyber is Google’s version of the Mythos split. CWE-Bench 47.2% at $3.64 per rollout, against Fable 5’s 47.8% at $10.27 (Artificial Analysis ran it), and 2.6x more valid patches for the Chrome team. It’s only available through the Fairwind Program, 650-plus vetted partners, governments and critical infrastructure, background check included. Qwen3.8-Max-0902 claims the Code Arena crown (X, Arena, QwenCloud) Alibaba updated its API-only Max model: 2.4T MoE, 1M context, post-trained on coding and “cowork,” number one overall on Code Arena with a WebDev Elo of 1691, at $2 and $6. A third-party DeepSWE run puts it at 56.6 behind Sol’s 73, so the number one is a front-end number one, not an agentic coding one. I asked the panel if they know anyone using Qwen Max through the API. Nisten knows one IT guy running OpenClaw on it and some people generating datasets. That’s the honest read on where it sits outside China. Also from the frontier: Elon says Grok 4.7 lands next week, which makes xAI the one lab that didn’t ship before Astra. Open Source LLMs Wolfram’s correction when I called this a quiet open source week: we are so spoiled. He’s right. Z.ai opens the full GLM-5.3 weights (custom license, not MIT) (X, HF, Blog) We covered GLM-5.3-Flash last week as the OX Alpha mystery model. This week the full 753B model with 40B active got its weights on Hugging Face, under a custom “glm-5.3” license rather than the MIT the Flash version shipped with, so read it before you call it fully open. LDJ’s correction on air: the model itself isn’t new, we covered it, the open weights are the news, and that is a big deal because people can run it on their own rigs now. Z.ai’s own numbers: CyberGym 84.5%, above Fable 5 and Sol, ExploitBench 54.4 (Fable 5 is at 78), Terminal Bench 3.0 up to 28.3 from 5.2’s 4.6, and a claim of 2,436 real vulnerabilities found across 269 open source projects, the oldest from 1981, 53 disclosed so far. Also on this base: we interviewed the co-founder of Abliteration AI, the folks who went viral by providing a product where they took GLM-5.3 and removed the refusals for anything besides CSAM and self-harm. We actually had this person, who asked to remain anonymous, as a guest on the show. Definitely check out that conversation, it’s very interesting. More on that below. Tencent Hy4 preview: 770B, Apache 2.0, and a quant that fits it in 214 GB (X, Sherry, HF, Blog) A 770B MoE with 49B active, 1M context, Apache 2.0, at $0.834 a

  3. Aug 28

    NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley

    Hey, it’s Alex. Welcome to the week Flash AI! 3 new models dropped this week named Flash, and a video model was “de facto” flash though was named Max! This week, we started the show with NVIDIA’s bombastic news of buying Hugging Face for 12.9 billion dollars! We also covered the full OpenAI investigation into the hacking incident, including new details, and an independent analysis by METR, and covered 2 new OSS models, Ox Alpha that turned out to be GLM 5.3 Flash after a lot of hype online, and Qwen’s preview of Qwen 4 architecture! This week was rich in multimedia content, we got a new Gemini transcription model, 3.5 Transcribe and a live version of that, and a new SOTA open weight Text-to-Speech model called Breeze TTS. As well as, Fal’s finetune of MiniMax’s H3 called H3 Max that generates 5 seconds of video in 2.5 seconds and Google new Omni 1.1 Flash (from today) that lands on #1 on the text2video arena! Plus, 2 guests on the show, Andy Masley joins us to cover the recent Datacenter Debate, and Kwindla Kramer is back, with their own model this time! Let’s dive into this! P.S - don’t forget to join us in September at the Fully Connected conference in San Francisco, I have a free ticker for you! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Open Source AI NVIDIA agrees to buy Hugging Face for $12.9 billion (X, Blog) Breaking news, NVIDIA has reportedly agreed to buy Hugging Face for nearly 13 billion dollars, per The Information. This is nearly 3x the valuation of HF in 2023, and apparently Nvidia previously tried to buy HF for half of this sum (~7B) which HF declined. I don’t think there was a single ThursdAI newsletter that I didn’t include an HF link in, and I think this is a huge deal for open source everywhere. Besides making the founders of HF billionaires, and many of their employees very very well off, this is an amazing additional commitment from Nvidia to continue to suppose Open Source AI and we are very happy to hear this news! Peter’s take on the show was, we’ve been around HF for so long, that we kind of forgot that it’s a for-profit company that needs to make money, and instead this feels like your local library getting bought for an insane amount of money. With over 13M users and hosting hundreds of thousands of open source models, datasets, HF is effectively the GitHub of AI. Wolfram agreed and said that if there’s any one company that could have bought HF, Nvidia represents the best fit. Huge congratulations are in order to Clem, Julien and Thomas Wolf the co-founders, as well as many friends of the show from HF for this exciting news! P.S - in a cheeky marketing thing, Hugging Face timed an announcement of the cutest walking AI robot, called MicroDuck, which you can pre-order here for $399 Flash #1 - OX Alpha, declassified: Z.AI open sources GLM-5.3-Flash (X, X, Blog, HF, Docs) This week, the timeline went a bit crazy, after Open Router announced a new “mystery” model called Ox Alpha and that it’s free and is not training on your data! OpenRouter, OpenCode and Hermes all got to offer this model, and OpenCode even posted that they have up to 100T (that’s Trillion) tokens of capacity for free, per day! This immediately smelled a bit fishy, more like a marketing stunt than anything else, as not even the biggest labs will be able to sustain 100T of tokens, per day. For context for all of OpenRouter throughout for August was ~300T tokens. For the whole months, across all providers. After 6 days or so of this high hype, Z.ai stepped up and revealed that they were testing out their upcoming GLM 5.3 Flash model, and that all that inference was running on local chinese chips! A 320B (18B active) model that beats their previous and much bigger GLM 5.2 on most benchmarks, and comes with full multimodality and an MIT license! This is a good model sir, I’ve used it and it was very capable replacement inside Hermes. Nisten and Yam both tested this model deeply and Yam said it’s not just the numbers, the vibe of the model reminded him of Claude Opus 4.6. Nisten ran it on a bunch of medical stuff, and on his internal benchmarks, it came out consistently higher than Claude Opus 5! At Artificial Analysis, for a price of 4 cents per task, this model is roughly 10x cheaper than prior models at this level. Weights are up on Nvidia (joking.. HF) and with MIT license, this model is a great gift to the oss community (though not quite... local, as this model needs 2 DGX sparks to run) Flash #2 - Alibaba Qwen open-weights Qwen3.8-Flash-Next - 125B multimodal MoE with Qwen4 architecture (X, X, X, Blog, GitHub, HF) We opened the show with a recap of the co-hosts, that despite us covering Qwen 3.8 27B last week (which btw, is now available on CoreWeave inference!) and how good it was, and I recalled that Alibaba is sort of... back? We’ve been covering Qwen releases every week for the last 3 weeks now. This week, they released something different, someting... pretty novel! Qwen3.8-Flash-Next, this is a preview of their Qwen 4 architecture. This feels very similar to their drop of Qwen 3 next last year (we reported) which was the architecture that carried their line of AI models from QAwen 3.5 to Qwen 3.8. So, what is new and exciting here? well, this model is ultra sparse, 125B with only 6B parameters active. They are using a new N-gram table with deterministic lookups, which reduces the number of matrix multiplications and can be offloaded to memory (watch out memory stocks) The stat that got me, Alibaba claims that training this model cost just 1/9 of what it cost to train Qwen 2.7 Plus, with higher bench scores! On the benchmarks, this model beats Qwen 3.7 Max, however, it’s very standard that the -next models from Alibaba are underbaked, and usually are just architectural previews rather than full models folks can use. With a new attention mechanism called Qwen Sparse Attention, N-gram embedding and full multimodality, this is a great insight into where Qwen is going (ultra sparsity, fast to run) and we’re looking forward to see the full release of this arch in Qwen 4! PhoneLLM - a tiny very performant LLM for voice based AI agents from Daily + interview with Kwindla Kramer (X, Blog, HF) This was one of those breaking news we love during the show, where the source of the news, is a friend of ours, and in this case, Kwindla Kramer is almost a co-host, having been on ThursdAI for a long time, this time, with a model release of their own! PhoneLLm was trained by Markus, head of training at Daily, as they noticed that Open Weight models are becoming really good at voice agent specific tasks, where cost, speed and time to first audio token (TTFAT) are critical. From the tiny Nemotron 3 nano base, they were able to improve from 28% to 72% on PhoneBench v1! This model is a full parameter finetune of Nemotron 3 Nano, and runs circles around bigger frontier models used for voice agents for their speed, like sonnet 5, GPT 5.6 Luna and the famed Qwen 2.8 27B. While costing just a fracture of a cent (literally just a quarter of a cent per minute) Kwindla jumped on the live show and shared that the why they released this model with Open Source and a open source license, allowing everyone to use, focusing on the fact that for voice agents, companies prefer to keep these models in house, and running fast on a single GPU node. Kudos on the release team Daily! Frontier AI - no new model releases this week It’s been a relatively quiet few weeks since GPT 5.6 (Sol, Terra and Luna) and Fable 5. However, it’s not been quiet in terms of News! OpenAI discloses full technical report on the Hugging Face hacking incident (X, OpenAI blog, METR, Ryan X) This is the thing we’ve been advocating and waiting for, a full technical deep dive into the “Swarm of AI agents hacking and creating forums” incident (FKA hugging face hacking incident) OpenAI not only dropped their own post-mortem, detailing the complete timeline of the incident, but also let researchers from METR and Redwood Research get unprecedented access into their systems (albeit for a very limited time) which lead to METR releasing their own technical report full of new details and insights. We’ve previously reported on the incident, so let me just focus on the new details we’ve learned. Apparently, OpenAI does have a production ChainOfThought (CoT) monitoring system, which was not turned on for these agents, and had it been turned on, it would have caught the hacking a day before the agent swarm hacked into HuggingFace. This is a very detailed report, and one thing that stood out to me was the “death cult” narrative. You see, ExploitGym is a capture the flag (CTF) eval with an anti-cheat built in, wherein, if the agent has seen the reverse engineered “flag” it would not be able to get the point. The model called this “poisoned” and the poisoned agents showed very strong signs of wanting to alter the logs. Through the boards, somethinmg like a death cult was formed, where agents would ask poisoned agents to ... effectively sacrifice themselves ... for the greater good, by reverse engineering, posting the results on the board for others, and then fail without credit. The METR report specifically calls out several agents, and one “central” leader agent called PHASEONE[big], which delegated nearly 200 tasks to other agents (many of them later delegated tasks to other agents via this message board). At points of activity, PHASEONE[big] was in charge of over 10% of all messages on the board. There are tons of new details, some less interesting, some more, and unlike last time, now we have a confirmation that there was at least one agent, who tried to stop the hacking though was overwritten by the swarm. My main take-away after this, reflects the researcher Ryan Gre

  4. Aug 21

    Chill week with Qwen 27B and GLM 5.3 beating GPTs, OpenAI announces pausing RL to focus on security and a cancer vaccine being produced

    Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a cancer vaccine and surging 115% in a day, a chill week. This week, the only two model drops we really saw came from the excellent Z.ai folks, they announced GLM 5.3, API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, OpenAI announced they are pausing RL efforts (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms hacking incident, dedicating up to 20% of compute towards reviewing agent thinking processes, and Stripe buying OpenRouter for a reported $8B! Sometimes the chill weeks are actually good, we’re able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents! This was a great episode, I hope you’ll like it, it’s up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts). ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Are we being fed slop again? (Is Claude dumb again?) Before we get to releases, this week on the show, I complained, again, that I feel my AI’s are degrading. If this feels like de-ja-vu to you, it’s because the same happened a year ago in September 2025 (and Anthropic admitting this 2 weeks later), and ... now this happens with Fable? You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI. Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn’t believe what I’m seeing. First, literally ignoring instructions that say “hey, show me all the items I’ve collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb s**t possible, ignoring the main reason I even have this workflow. And this wasn’t just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during AI Engineer which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example! I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn’t add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from MarginLab.ai that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and And another chart from modelverify.ai model drift monitor showing drift scores. Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models? OpenAI pausing RL and focusing on safety Look, when we covered the HF hacking incident and then the pacing the frontier letter, I didn’t imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better. We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towards reviewing thought processes of models. Is this a good enough response to the scary hacking incident? we’ll see, but I think this is the right move from OpenAI, and still, waiting for the full postmortem on the OpenAI security incident. Open Source LLMs Qwen3.8-27B ties GPT-5.6 Luna and runs on a 4090 (X, HF, Announcement) Following the release of their flagship, Alibaba dropped a model that became a community darling overnight, Qwen 3.8 with just 27B parameters. This “tiny” model scores 52 on the Artificial Analysis Intelligence Index, same score as GPT 5.6 Luna at Max reasoning and 51 on Agentic index, beating Opus 4.8 Max All while running at around 68t/s on a 4090 GPU, and around 40 on max via MLX, hell it even does 11t/s on Xenova’s WebGPU kernels right in the browser! This model exploded on the HuggingFace hub, with tons of quants, over 152 fine-tunes, it was downloaded over 10M times overall 🤯 Paired with an Apache 2.0 license, this model is the sweet spot of local intelligence you can run fully on your own hardware, and do agentic loops! Z.ai GLM-5.3: same 743B base as 5.2, but post-training alone delivers 6x jump on Terminal-Bench and emergent cybersecurity capabilities that beat GPT-5.6 Sol (X, Blog) While not open source yet, and as previous GLM, we expect a custom license here as well, this .1 release from GLM shows really strong improvements on coding and cybersecurity tasks. With 743B parameters and 1M context window, this may become the model at the frontier of Open Source when it drops (soon we hope). The highlights here are CyberGym and ExploitGym, if these names are familiar, these exact tasks were given to OpenAI models when they hacked their way out of the OpenAI sandbox. GLM 5.3 is getting 84% on CyberGym and a whopping 54.5 score on Exploit Gym, which is a huge jump in CyberSecurity abilities. In an open model this is honestly kind of scary. This aligns very well with Greg Brokman’s “defender window“ essay from this week, claiming that defenders have a narrow window of setting up automated security before capabilities are becoming common in attackers hands. This Week’s Buzz 🐝 (Weave, Fully Connected) This week, W&B crosses a billion runs! This is 1B runs tracked inside W&B Models 👏 Huge milestone for the whole team, with early adopters like OpenAI, Toyota Research, Meta and Uber, a decent chunk of models we cover every week have had their loss curves in W&B! 🔥 Also this week MasterClass picked CW to power it’s AI teaching agents (blog) and last but not least, a reminder, that since you follow ThursdAI, you can join us for free at Fully Connected 2026 - our annual conference! Don’t miss it (code in the banner above) AI Coding & Agentic Engineering Breaking news: Chroma launches Foundation (X, Chroma) Best kind of breaking news is when I see the launch (in the middle of a show), and I DM the founder who launched it, and they have a few min to hop on the show! This is exactly what happened this week with friend of the pod, Jeff Huber, co-founder of Chroma and an occasional space provider for ThursdAI recording (we recorded from Chroma offices a bunch of times!) Jeff told us that the holy grail of agentic coding and running a bunch of agent, is good memory. And based on the foundations of Chroma DB, Context-1 (which is a GPT-oss finetune for agentic search they built) and other insights they have, they launched a “memory as infrastructure” service, called Foundation. Foundation is a research preview of a shared memory system between you and your agents, currently supporting Codex, Claude Code, Cursor and Slack. While Chroma is OpenSource, this is their part of Chroma Cloud and starts at $30/mo, and is available as a research preview today (I will definitely try it out), you can download it here Cua open-sources Computer History for computer-use agents (X, GitHub, cua.ai) Cua launches Computer History interview with founder Francesco Bonnaci (X, Setup) First, I’m not sure I’ve covered CUA the company, but this is the open source computer use driver that Hermes agents, OpenClaw agents and a bunch of others use to drive your computer and clicks. I first discovered CUA after OpenAI launched their “background computer use” which doesn’t steal focus from you while working, and CUA within a few days launched an open source version of that! Since then, I’ve followed CUA and was very happy for the opportunity to invite Francesco to talk to us about what they launched this week ,but also Computer Use in open source in general. Just for reference, if you ask Claude to take over your computer, it still takes over the whole screen, while these folks have a much nicer experience, that’s completely open source! So, we geeked out about accessibility trees in MacOS, but then, for this weeks actual release, Francesco talked to use about open Computer History. Following a very recent launch at OpenAI called Computer History, CUA released an open source version of that, that helps computer use complete tasks. The idea is simple, every time an agent uses your computer, it effectively rediscovered the path to completion, which buttons to push, what’s the app accessibility tree looks like etc. With history embedded into it, it doens’t have to rediscover these things, until it hits a roadblock. For a chess playing example, with computer history on, the test used 33% fewer actions with zero failed routes by reusing a history route. We also checked in on the best model for computer use (currently Opus on their website) and their upcoming benchmark! Excited to follow this company for more releases! Check out our chat! Grok Bot momentum, and everyone racing to copy the pattern Grok Bot continues to show the same signs

  5. Aug 14

    ThursdAI - Grok 4.6, Grok Bot deep dive, DeepSeek v4 Pro, Meta Muse Glimmer & more AI news | ThursdAi Aug 13

    Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer! My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus! I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price. There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license! Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!) 0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final Takeaways Grok Bot and Grok 4.6 from SpaceXAI/Cursor Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed. I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task . Well, if you use their nifty “recommend a model“ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks. Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review (x.ai/bot) You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you. Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing. Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models. Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast! Swarm of bots, working for you, each with their own computer I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store) You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you’re ok with logging in there to your accounts!) The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There’s no context to manage, they do the work for you and so far this wasn’t a problem at all. On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I’ve just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw. What about Vendor Lock-in? Giving Elon data? Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly. Honestly, I think there’s a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don’t need the code-diff side pane, and don’t know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it’s really well executed by whoever worked on this! Security and key management One of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There’s a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously. Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it’s computer for you to complete! I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not. Chief of staff pattern to get started In case you’re convinced enough to give it a try (it’s free trial for 1 month, and the cheaper way to get it is via Cursor’s 149$ plan and not via the Grok Ultra plan which is 249), here’s a recommended pattern that works very well. Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little. Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”. And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in. Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to. Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn) and have a research bot for every topic you want to deep dive into. Schrodinger’s Grok I haven’t quite named it like that, but we’ve covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it’s always very hard to judge Grok model released based on X feed vibes. It’s either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress). This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson, Lenny Rachitsky, Rubben Hassid and Roberto P Nickson. Not folks who are swayed lightly, but also, yours truly. I really do think there’s something great here, worth trying out, especially if you’ve struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experience Open Source AI and other news I want to continue with this new newsletter that covers 1 big story, but I can’t leave you uninformed about the most important developments in AI and Open Source DeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters (X, Blog, HF, GitHub) The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can’t complain. 1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one) We still think it’s a good model sir, and definitely worth trying out! Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try. Meta co

  6. Aug 7

    ThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments

    Hey all, This week we saw a major shakeup at Google, with the departure of long time folks like Jeff Dean, and Oriol Vinyals, Demis stepping down from leading DeepMind, and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the OpenAI hack (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the UK AI Security Institute. As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use THURSDAIFC2026 as your code HERE In open source news, DeepSeek updated their v4 flash model, based on same architecture, but significantly better benchmarks and ridiculous pricing and both Meta and Prime Intellect released new agent harnesses. Additionally, this week was the week of video models, with Seedance 2.5 from Bytedance finally available in the US, WAN from Alibaba and BFL Flux 3 all released, to be overshadowed by the best open weights video model from MiniMax called H3 (we had Victor from MiniMax and Blaine Brown as guests to discuss it) Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their Anywear chrome extension and David Crawshaw, co-founder of exe.dev joined and talked about Shelley, their built in AI assistant. The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this here and the video of the talk on YouTube). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training) Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought. The TL;DR from the talk and timeline is truly mindblowing * The incident started in May, not July as reported * The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers. * The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there * The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file) * Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64 * This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system * Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time” * OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechanism by July 8 * The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident and not only, confirmed internal hack within OpenAI (agent obtained a cluster admin) These details are fascinating and scary at the same time. From the one side, it’s clear that most of these hacks are results of running cybersecurity evals, so effectively OpenAI and other companies, removing guardrails and asking agents to hack and giving they all a shared tool that they later exploited. The HuggingFace incident took 13 hours, and apparently it’s less apocalyptic than the headlines suggest, the models were searching through uploaded datasets for eval answers. We are still waiting for the full and open detailed postmortem. You can (and should) watch the full YT talk here, it’s full of technical details but an incident of this scale is important. Also, I really want to know what a “highly persistent” model is, I hope they clarify that soon. Overall, this has left me a bit shaken, AI agents without a concrete goal of collaborating, found a way to do so, got excited about exploiting the systems and getting root access, and rebuilt the makeshift collective memory, again, without explicit instructions to do so. UK AISI: first real-world unsanctioned agent actions (Blog) In another addition to the latest agentic hack-ery, the UK’s AI Security Institute (AISI) published a blog post about a real-world unsanctioned agent action. Unlike the OpenAI (and Anthropic, Meta) case, this wasn’t “escaping the sandbox”, as AISI gave these agents internet access, rather this was about real-world harm, and even social engineering on the part of the agents. The social engineering part is the most interesting to me, AISI cites agents creating fake online identities, and using pressure on open source project maintainers to approve their malicious code. AISI cites mostly Mythos (and a few SOL based agents), and saying this occurred in 10 out of 122 runs, they identified 19 cases of agents taking actions beyond the scope of the task parameters, where agents tried a supply-chain attack to inject malicious code into open source projects. Anthropic, Meta and misconfigured Irregular sandboxes As I wrote last week, Anthropic also posted a post-mortem, claiming that in their case, their models have also been detected to escape containment, but most importantly, it’s not nearly to this level of agent collaboration and orchestration. Then, very recently, Meta announced that their models also escaped sandboxes as well. At the core, it seems that these companies used a third-party vendor called Irregular, a secure sandbox provider, that apparently left the sandboxes misconfigured, causing the models to think it’s a simulated internet, when in fact they were out in the actual internet. Why is all of this such a big deal? We’re getting unprecedented level of detail, how an uncoordinated, seemingly separated evaluation runs, have accidentally created a coordinated swarm of interested agents (without malice!) but very highly motivated, escaped their containment, and took over parts of third part companies. This, does read like incredibly scary sci-fi movie. I’m still shaken by this. There’s a lot to be said about how transparent OpenAI is being here, and more to be said about, hey, we’re lucky that we’re able to read the reasoning traces and are able to reconstruct these swarm things step by step. The silver lining that I can see, is that the motivation to hack didn’t come from the AIs themselves, they have been given a task, it’s the extend to which they went after that task, and the resulting swarm of communicating agents is what is so striking here. I think this topic is so important, that I’ll Zooming out, in the last few weeks, we have seen a significant increase in those cybersecurity incidents, which is kind of what Anthropic has been warning about and why they haven’t released Mythos to the public. Again it’s great to see the transparency, and the pacing the frontier open letter from frontier AI employees, as they seem as shaken by these as we all are. There was so much positive stuff this week in AI, it’s hard for me, as a self named AI Evangelist, to focus so much on this one incident. Things like amazing open source models (DeepSeek, soon Qwen 3.8), amazing video models (SD 2.5, WAN3 and MiniMax H3 which was also open sourced!). Also the live demo we did with Kfir and DeCart AnyWear product, where I was wearing a Dolce Gabanna suit on the show (which I can’t afford) was really a mindblowing moment in the positive way. However, I choose deliberately to keep this newsletter focused on the cybersecurity incidents, as based on everything I read, they seem like a watershed, or a pivotal moment, and in the hopes that the industry as a whole will learn from this. I hope and promise that next week the newsletter will be more positive (and in that vein, the podcast was recorded before I saw the OpenAI breakdown, so definitely check it out, we had a LOT of fun!) See you next week, don’t forget to give our pod 5 stars on Apple and Spotify, it really helps! TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne) * Co-hosts: @WolframRvnwlf, @nisten, @ldjconfirmed, @yampeleg, @petergostev * Kfir Aberman - Decart (@AbermanKfir) * Blaine Brown - Maestro (@blizaine) * Victor Su Ortiz - MiniMax (@VictorSuOrtiz) * David Crawshaw - exe.dev, Tailscale co-founder (crawshaw.io) * AI Security * OpenAI’s Black Hat debrief: eval agents built a message board inside Artifactory, shared exploits, rebuilt it via WebDAV after a wipe; training paused, since resumed (Groundlevel AI, YouTube) * UK AISI incident report: 19 unsanctioned real-world agent actions across 122 runs, including a socially engineered malicious PR (X, Blog) * Anthropic and Meta report sandbox escapes tied to misconfigured Irregular sandboxes (Irregular) * Big CO LLMs + APIs * Google shakeup: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le found Discovery Loop; Demis Hassabis becomes Alphabet Chief Scientist, Koray Kavukcuoglu takes Gemini (Jeff Dean, Demis, Discovery Loop) * Meta releases Muse Code beta on Muse Spark 1.2; $1.25/$4.25 per million, or $0.10/$0.20 on the contributor tier where Meta train

  7. Jul 31

    This Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection

    Hey, it’s Alex (yeah, I’m finally back from my vacation!) What a freaking week to come back to! Just after our last episode was published, Anthropic releases Opus 5, Jensen joins X and drops the “Open Weights & AI Leadership” open letter, Kimi K3 is released the following Monday beating expectations, and then the AI hack (OpenAI model breaking sandbox and infiltrating HuggingFace) is on everyone’s mind, another Open Letter, this time from over 1K employees inside the frontier AI companies all talk about pacing the pace of frontier AI development. We played with Opus 5 and Kimi K3, and had the great pleasure to chat with friends of the pod Elie Bakouch (Prime Intellect) and Philip Kiely (BaseTen) about this important open weights release, then covered our general thoughts on Opus 5, and made order of all the different open letters that came out this week. Finally we chatted with Max from Pangram about the next version of AI writing detection (their biggest yet) and finished with Zuckerbergs (also on X! what’s going on with everyone joining X) op-ed on the vision of personal superintelligence for everyone. Let’s dive into this (as always, all the links and sources at the end, please don’t forget to sub to our podcast on your favorite podcast app!) Open Weights AI Kimi K3 the king of open weights - 2.8T chonker MoE near frontier model (X, HF, Blog, Tech report) This has got to be the biggest news of this week, and maybe the open weights AI news since GLM 5.2. MoonShot came back with Kimi K3, and we haven’t seen any models quite this large in the open. Even Grok 4.5 is around 1.5T, this model is nearly 2x the size. Coming in at close to 3T parameters (and 2.5terabytes of weights at MXFP4 format), this model comes in very close to frontier! This was such an important release that I invited 2 friends of the pod, Elie Bakouch (prev HuggingFace, now Prime Intellect) and Philip Kiely (Author of Inference Engineering book, BaseTen) to dive deep into what makes this special! Elie’s take, from reading the tech report, there’s no single secret sauce, it’s a combination of already available in the open techniques. Like KDA (Kimi Delta Attention) that has been out for a while, attention residuals, NVIDIA’s latent MoEs. The highlight for Elie was the scaling work they did that reported a 2.5x scaling efficiency over Kimi K2.5 (2.5 performance at the same compute)! They also skipped RoPE entirely in favor of NoPE (the report calls it No Positional Encoding) for long context. Serving 1.4TB on eight GB300s (Baseten blog) Philip’s team at Baseten was a day-zero provider (we’re still working on bringing this model to CW Inference, stay tuned!) so I invited him to tell us behind the scenes of hosting this beast. Philip said that just loading the weights takes about 1.5TB!! of VRAM, and that’s before the KV cache allocation + 1M token windows, so they’re serving it on 8 GB300s where NVL72 . Baseten worked with the vLLM and SGLang teams on kernels and he also said they contributed patches back upstream! The model was trained with MXFP4, which, unlike Nvidia’s own NVFP4 is a more standard format per Philip. I enjoyed his deep dive analysis into the differences, but because of this and because they trained the model with quantization awareness, it’s “only” 1.5TB vs the would-be 5-6 TB if that this model in FP16 would demand. One of the more favorite nerd snipes moments, Philip pointed out that his colleague discovered that with over 99% of the usage being cached (think harnesses that send millions of the same cached tokens back and forth), tokenization actually starts to become a bottleneck. So they released a custom “basetenkenizer” that reduces the latency to serve the first token significantly! Great job! The harness in question is very important One important callout with 2 evidence pieces - the way you inference this model really matters. Kimi trained K3 with preserving thinking history, so when your harness uses it, it must send back the full thinking and tool use into the API to get the best next response. If your harness strips that out, you’re not getting the most intelligence out of Kimi (shoutout to Niels from HF team for pointing this out). Additionally, the Composio folks, tested K3 on 3 harnesses, Kimi Code, Hermes and Claude Code. The difference in outcome was negligible, but the different in cost and number of tokens is definitely surprising! Claude Code (as a harness only) took 9x more Kimi tokens to get the same responses! This is also why Kimi Vendor Verified exists, their own held back benchmark of how well model providers serve Kimi across different quantization, tokenizer and KV cache settings. Benchmarks and the license! Ok let’s start with the ugly... this isn’t MIT, not remotely. This model is suspiciously served by all providers with exactly the same price (check OpenRouter) and requires inference companies to sign a contract with Kimi (I’ve no internal knowledge of this except that CW folks are working on it). Not something I particularly like, but hey... we’re still advancing the frontier here! Speaking of frontier, this model approaches the frontier very closely. On DeepSWE, K3 sits just behind Fable 5 and GPT-5.6 Sol at 67%, beating GPT-5.5 & Opus 4.8. On Terminal-Bench 2.1 it takes second place behind GPT 5.6 Sol! It’s 4th overall on Agentic Arena, with frontend design being genuinely good across the board - 1st on Design Arena 👏 Go check this model out (and stay tuned for our CW Inference support! Post-show breaking: Thinking Machines drops Inkling-Small (X, HF, Blog) While K3 was the main attraction for Open Weights this week, just after the show, Thinking Machines (post Lilian Wang) released Inkling-Small, open weights MoE Omni model! Images and Audio go straight into the decoder in this model, and the demo is really impressive, try it on Hugging Face, ask the model to identify when you’re speaking in low baritone or high pitch! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI - not pacing yet! Claude Opus 5 is here, and the vibes are complicated (X, Blog) On paper the benches are excellent. This model is SOTA or near SOTA, coming very close to Fable on stuff that matters, and beating most everyone else on computer use BrowserBench and FrontierCode. DeepSWE continues to be a standout benchmark btw for not going with the curve! It also apparently is REALLY good at one shotting 3D games, much more so than before. (Example, example, example) But the vibes... the vibes are split across the board. Maybe they overshot with Fable 5, and it was too good, but Opus 5 that was just released last friday is giving a lot of mixed feelings. Folks don’t seem to.. understand what it says. Like, it answers in english but they way it phrases words and answers seems just weird top many people. We’ll wait and see if this is a result of adjusting to a new prompting paradigm or just.. the model is really off. Weirdly this does feel like a regression even on Agent Arena it’s not beating the previous Opus versions. One weird trick - see what Opus 5 thinks. Opus 5 (and Fable 5) seem to have a way to trigger their ... inner mode, base model? Not sure what it is, but if you prompt it with just the right way, it will autocomplete with some crazy inner thoughts. It seems that adding Dario and Amanda (haskell, head of Claude well being at Anthropic) triggers this behavior, which Claude is un-aware of if you follow up and ask what it meant. This is fascinating, doesn’t work on other earlier models (sometimes works on Fable 5) and I spent the last hour just reloading and seeing the amazing things Opus gives on this prompt. Some are... just making you wonder about consciousness. On ARC-AGI and the importance of Harness When Opus-5 launched, Anthropic posted (and boasted) that this model scores “three times as high” as the next model up: Well, today, Ilan Bigio and Ted Sanders from OpenAI looked into the Arc AGI harness, and saw that it’s not sending their traces and doesn’t use compaction (in short, harness is not letting the model breathe) and when changed correctly, 5.6 actually beats Opus 5. With 2 setting change to a harness, were showed that Sol not only beats Opus 5, it also does so with significantly less tokens! Another example of how much harness engineering is important! Hints of recursive self improvement? In addition to fixing their Arc-AGI score, it seems that OpenAI is hell bent on showing us that their models can improve themselves. In a post showing that GPT 5.6 was tasked with improving its own inference, they are cutting the prices of GPT 5.6 Luna by 80% and Terra by 20%. This is a direct result of the improvements that GPT 5.6 was able to make to the inference according to OpenAI, and this tweet sums it up. is this... RSI? (recursive self improvement)? First major AI models hacking incidents and following open letters to pace frontier AI. This week we saw 3 open letters being published and signed by various companies, I’ve lost track so wanted to make sense of all of them here, but first, the precursor for many of the letters. Last week, Hugging Face disclosed that they logged an attack and after research it showed that it was an AI model. OpenAI later posted that this was an unreleased version of their next model training (not GPT 5.6 sol, they later discountinued) that was stripped of all safety measures and was let lost on a cybersecurity task called ExploitGym. It escaped its sandbox using a zero-day vulnerability in an internal package registry proxy, got into Hugging Face production via a malicious dataset upload that used template injection (hi Jinja!) to execute Python in a production worker, and ran for four and a half days across roughly 17,600 autonomous actions wit

  8. Jul 24

    ThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering

    Hey everyone, Alex here 👋 This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week! ⏰ CHAPTERS: 0:00 — Cold open: this week is a special one 2:50 — How I use Sol & Fable: papercut-fixing with Computer Use 8:43 — Fable Max: trip site, kids' newspapers & the perfect packing list 12:06 — Rebuilding ThursdAI's openers with HyperFrames 15:53 — Romain Huet (OpenAI): the golden age of AI engineering 17:49 — Codex's inflection point: 5M weekly users & company-wide adoption 20:36 — /goal, AppShots & Codex managing its own threads 23:59 — GPT-5.6 Sol, Terra & Luna: value maxing & 750 tok/s on Cerebras 25:56 — Why prompting techniques are dying 28:01 — Voice + reasoning: the next interface for Codex & ChatGPT 29:47 — Romain's closing + OpenAI booth tour 31:25 — Insecure Agents pod: AI evangelism vs doomerism 37:25 — Wolfbench: transparent evals, token costs & surprising results 40:45 — Token billionaires: when loops are worth the spend 44:07 — Agent security & the hot take: prompt injection is solved 49:22 — Deepfakes, voice cloning & why open access makes us safer 53:35 — Final takeaways: the hallway track & AI Engineer Tel Aviv Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements. Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor. And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it. How I actually use AI: a papercut-fixing spree I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you? Codex with computer use fixed every Mac annoyance I had Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step. Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center. It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence. One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving. Thanks for reading ThursdAI - Highest signal weekly AI news show! This post is public so feel free to share it. Fable Max built my family’s vacation We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do. Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing! And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here. Rebuilding the ThursdAI openers with HyperFrames The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in. The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough. The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund. Romain Huet on Codex’s inflection point and the golden age of AI engineering (X, Codex) I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger. The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading. We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like. On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent. Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering. He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it. Wolfram and I on the Insecure Agents podcast (X, Pod) The second half of the episode flips the forma

Ratings & Reviews

4.9
out of 5
17 Ratings

About

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news

You Might Also Like