Latent Space: The AI Engineer Podcast

Latent.Space

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space Sponsorship and business inquiries: business@latent.space www.latent.space

  1. 4d ago

    Academia is for Ambition — Alex Zhang, MIT

    Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks! While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning. This year we are proud to feature the work of Alex Zhang of MIT. From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. RLMs took over the timeline early this year: and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra: and is even today, influencing new research that has more extreme implications than RLMs: We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie. We discuss: * Why AI-generated GPU kernels still leave substantial room for human expertise * How one expert insight can potentially replace enormous amounts of brute-force token search * Why PhD students should take research bets that initially look trivial, weird, or pointless * What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste * GEV and why a language model does not have to mean an autoregressive text-to-text decoder * Why Claude Code, Codex, and Pi are structurally more similar than they look * How harness design can improve compositional generalization across tasks and domains * RLMs: context offloading, code execution, recursive subagents, and shared memory * Prime Agent, continual harnesses, and persistent agent-to-agent communication * Why the model you query in the future may secretly be an entire swarm or scaffold * OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving * Why much of an agent swarm may be wasted search — and why convergence is still hard * Kimi versus OpenAI and different approaches to multi-agent systems * Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work * Why current frontier models may already have a large capability overhang * Speculative programmatic tool calling and overlapping tool execution with generation * Whether English, code, or an entirely new “Neuralese” constrains how models reason * AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet on Alex Zhang * Website: alexzhang13.github.io * X: @a1zhang Timestamps 00:00:00 Introduction 00:00:49 GPU Mode, KernelBench, and AI-Written Kernels 00:07:38 Human Expertise vs. Brute-Force AI Search 00:13:20 Research Taste and Taking Big Bets 00:19:28 GEV and Rethinking the Language Model 00:29:03 Video Game Agents and the Harness Problem 00:31:01 Why Claude Code, Codex, and Pi Are So Similar 00:36:42 Harnesses as Compositional Generalizers 00:44:24 RLMs Explained 00:52:01 Prime Agent and Persistent Subagents 00:57:41 RLMs in the Wild 01:00:30 OpenAI Swarms and the Future of Language Models 01:07:26 Open-Endedness and Sakana AI 01:15:52 Kimi vs. OpenAI Agent Swarms 01:20:06 Capability Overhang and Speculative Tool Calling 01:28:19 Neuralese, Future Research, and AI for Science Transcript Introduction: Alex Zhang, RLMs, and GPU Mode Swyx [00:00:00]: All right, we’re here in the studio with Alex Zhang, I guess most famously of, RLMs, but you have a few other affiliations. Welcome to the show. Alex Zhang [00:00:12]: Yeah. Thank you for having me. Swyx [00:00:13]: Yeah. I guess GPU Mode as well? Alex Zhang [00:00:15]: Yes, GPU Mode as well. Swyx [00:00:16]: You were shepherded in by Mark Saroufim. Not everyone gets that kind of welcome. Alex Zhang [00:00:19]: Yep. Yeah. Yeah. I’m very close to all the people in GPU Mode, so yeah Swyx [00:00:23]: Yeah Alex Zhang [00:00:23]: We often end up working together in various capacities, like even beyond just GPU Mode itself, so. Swyx [00:00:29]: Yeah. Can we explain, so people who are not that close Don’t know about this. It’s just a-- it’s a Discord. It used to be focused on, I guess, CUDA Mode, and then generalized a little bit. it was started by Mark. Alex Zhang [00:00:41]: Yep. Swyx [00:00:42]: It was basically like. To me, it’s like the hiring pipeline of the PyTorch team. Swyx [00:00:45]: And then you left PyTorch. Alex Zhang [00:00:47]: Yep. Alex Zhang [00:00:49]: Yeah. So it used to be, I think, it started actually around when I was in college, in like 2023. I think it was started by Mark, Andreas, and Jeremy Howard. The original premise was just like, it was a GPU-- or it was a Discord dedicated to learning how to write GPU kernels, and they had, like, lectures. That was basically the extent of it. and I got interested in it because I was writing GPU kernels. It was actually out of, like, pure chance. I was interning at Snapchat at the time, and From CUDA Mode to GPU Mode Swyx [00:01:22]: Rexis. Alex Zhang [00:01:23]: Yeah, I was very bored with Rexis. So, they had a project where, like, they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper. Swyx [00:01:35]: Yes, we’ve covered it on Paper Club. Alex Zhang [00:01:36]: Yes, yeah. So I was interested in whether or not you could write specialized kernels for it at Snapchat. it didn’t. Nothing really came of it, but I joined GPU Mode. At the time, it was called CUDA Mode, I think for, like, legal reasons or something, they changed the name. But I met Mark, I met Matei, I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn, which was now what you see as the leaderboard today. But the general idea was like. I think all of us had this, like, intuition that GPU programming is, like, very similar to if you guys have done, like, competitive programming. It’s a, it’s a not. I don’t mean to say, like, they’re transferable skills. Popcorn, KernelBench, and Automating GPU Kernels Swyx [00:02:17]: You have constraints. You code golf a little bit. Alex Zhang [00:02:19]: Yep. Swyx [00:02:19]: Yeah. Alex Zhang [00:02:19]: Yeah. And there’s like. There’s actually a surprisingly small space of optimizations that people do. and there’s actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that, like, if you had enough data, like, in the same way that Codeforces, there’s millions of problems. If you could do this with GPU code, like, you could scale and automate kind of GPU kernel development, which for researchers is a huge deal. ‘Cause I think one of the bigger bottlenecks. Like, if you look at, like, Mamba, for example, like, they release the paper with kernels because otherwise, like, you can’t really use it in any meaningful way? And not everyone has, like, a Tri Dao on their team. So we’re very interested in this. KernelBench kind of spawned from that too, of like, can we get LLMs to automate, GPU kernel code? And I think that was like a. It was a very fun time. It was like between college and my PhD, and yeah, I had a really pleasant time doing stuff with GPU Mode. Now I kind of just help with the lectures sometimes. I’m not as involved, and I think in general, like, we don’t have as many competitions as we used to. But, yeah, I still keep in touch a lot with everyone there. Swyx [00:03:27]: Is there a friendly rivalry? Because I think the previous community that used to do this was like MLSys, MLPerf Alex Zhang [00:03:32]: Yeah Swyx [00:03:32]: Kind of thing. Is there a friendly rivalry? Is this like just new generation MLPerf, or what’s going on? Alex Zhang [00:03:38]: The nice thing about GPU Mode is that it is also a community in the sense that, like, a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that. And, like, the competitions are, like, somewhat, not secondary, but, like, you can participate in them to learn. I think with a lot of, like, MLSys, MLPerf kind of benchmarks, like, for the most part, like, only serious labs and companies participate. Like, seriously in them, at least. That was my understanding of it. I could be wrong. But I think also beyond GPU Mode now, one thing that has been really exciting is there’s a lot more websites and, like, people that work on hosting competitions. Like, I think there’s this. Alex Zhang [00:04:21]: I think there’s this website called, like, LeetGPU or something, and it’s like leet code for GPU problems. Swyx [00:04:26]: Wow. Alex Zhang [00:04:26]: There’s, like, other ones that I’ve. Like, we’ve, we’ve seen. Like, there’s many that have kind of spawned and, like, talked on GPU Mode, and like, it’s very exciting in the sense that I think GPU programming used to be super niche, like when I was interested in it. And the only reason I got interested in it was Tri Dao gave a talk at Princeton because he was, applying for faculty. he is faculty there now, but I listened to his talk on FlashAttention in li

  2. 5d ago

    Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

    Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss: Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software. OpenAI clones Jev In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API: And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns. We discuss: * Why OpenAI thinks Computer Use has changed dramatically in just the last few months * Dots and what changes when every agent gets its own Linux computer * Why Computer Use can now complete some tasks faster than the average human * The path from human-level to “literally superhuman” computer use * Why modern agents are much better at debugging and recovering from failure * How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together * App Shots and why they give models much richer context than ordinary screenshots * Why Computer Use can close the loop between writing software and testing it * Trust, permissions, and safety when agents can make payments and operate websites * Async function calling and why models no longer need to stop reasoning while tools run * Mid-turn steering, WebSockets, and the architecture behind more responsive agents * UltraFast inference and how OpenAI is pushing frontier models toward much lower latency * The rapid internal story behind the Decisions API * Why Decisions API is more than structured outputs at low latency * GPT Live, fast tool calling, and real-time computer control * How OpenAI is already using Decisions API for support classification and internal workflows * Longer prompt caching, cache pre-warming, and cache-aware applications * Server-side compaction vs manual compaction for long-running agent threads * What should live inside an Agents API versus a developer’s own harness * OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs Ari Weinstein * Product & Engineering, Computer Use at OpenAI * X: https://x.com/AriX * LinkedIn: https://www.linkedin.com/in/weinsteinari/ Nikunj Handa * Product, API at OpenAI * X: https://x.com/nikunjhanda * LinkedIn: https://www.linkedin.com/in/nikunjhanda/ Timestamps 00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API 00:02:52 Dots and Personal Cloud Computers 00:04:59 Why Computer Use Is “180 Degrees Different” 00:06:04 From Sky to Self-Debugging Computer Use Agents 00:09:24 How Computer Use Sees and Operates Software 00:12:09 From Faster Than Humans to Superhuman Computer Use 00:16:03 Agents API: Trust, Permissions, and Safety 00:17:31 Computer Use for Coding, Testing, and QA 00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference 00:23:21 The Rapid Story Behind Decisions API 00:25:32 What Decisions API Is and How It Works 00:30:24 What OpenAI Is Building With the New APIs 00:32:23 Prompt Caching, Pre-Warming, and API Performance 00:35:20 Context Compaction for Long-Running Agents 00:37:13 Memory, Higher-Level APIs, and the AI Cloud Transcript Introduction: OpenAI DevDay and the New Agent Stack Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast Swyx [00:00:08]: We’re the first podcast after your livestream. Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today? Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it. Swyx [00:01:02]: And the API. Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote. Swyx [00:01:35]: And not to mention the Decisions API. Ari Weinstein [00:01:37]: Decisions API. Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models? Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate? Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast. Dots and Delegating Work to a Cloud Computer Swyx [00:02:07]: Yeah. Ari Weinstein [00:02:07]: They also make it a little bit less good at doing, like, long horizon, sort of sophisticated tasks. And so I think I would say it’s still an open area of research for how we, like, bring those approaches together. But, yeah, I’m really excited to see what people build with the Decisions API. Vibhu [00:02:24]: One of the interesting things is Dots now have attached personal computers. Ari Weinstein [00:02:28]: Yeah. Vibhu [00:02:28]: So it seems like they’re very much more persistent. You’ve been using them for a while. How should people push the bounds? Like, what should people aim for? What should they try? Personally, right now I use it for a lot of customer service. Like Ari Weinstein [00:02:41]: Cool Vibhu [00:02:41]: “Oh, this was wrong. I don’t wanna sign in. I don’t wanna authenticate.” Find whatever and just get it fixed. Ari Weinstein [00:02:45]: Yeah. Vibhu [00:02:46]: How should we push further? What should people try? Ari Weinstein [00:02:50]: Dots Are a really cool product because each Dot has access to its own Linux virtual computer in the cloud, which is different from our other products. you know, traditionally, we’ve have access to a browser in the cloud, or it has access to your own computer, but now you get your own entire Linux computer in the cloud. And so it can run full desktop applications, and it can also use a web browser. And so, yeah, you know, I think the powerful thing about Computer Use and the reason why I think it’s so, exciting is because it makes it so that the agent can do anything you as a, as a person can do, because all the software in the world was designed for humans, and now agents can use that same software, and you can delegate to the agent. So, yeah, like, anything that you would do on a computer, you can ask a Dot to do. Yeah, I think what particularly is useful is gonna really depend on who the end user is and what- what’s valuable in their life. but yeah, I would just start by thinking about, like, one of the things that you spend time on and how could you delegate those to an agent. Swyx [00:03:47]: Yeah, a lot of flight booking and shopping and honestly even, like, playing a game or whatever, right? Ari Weinstein [00:03:52]: Totally. Swyx [00:03:52]: Yeah. Ari Weinstein [00:03:53]: Yeah, I don’t know. For me, something I did recently, I’ve been working on. I’ve, subscribed to a meal prep service ‘cause I was trying to, like, eat healthy, you know? And I really like this meal prep service I found because it lets me customize the meals I order to, like, a high degree of granularity. So I can say, like, “I want this many grams of chicken and this many grams of rice.” but it was so complicated. It took me two hours to do an order, and I found that I could ask Computer Use to do it for me, and it did it in 15 minutes. so I actually saved two hours. it both did it eight times faster than I could, and it saved me two hours on GPT-6.1 Sol. Swyx [00:04:32]: Yeah. Ari Wei

  3. Sep 29

    Claude Code’s Next Era — Thariq Shihipar, Anthropic

    We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks! In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR: * June: Launched Claude Tag and Sonnet 5 and Fable 5 * July: Opus 5, /checkup. crossed $65B ARR * Last month: Fable/Mythos 5.1, and EFS (upcoming pod) * IPO target $2T, end 2026 ARR estimated $100B * Cowork/chat merged before did * Claude Mods * Dario endorses the same Pacing the Frontier message cosigned by all labs * Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects * Today: Sonnet 5.5! Today’s episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable: The Future of Mutable Software Pay special attention to Claude Mods (especially the cheatsheet): In general this is also the inverse of the other viral tweet from Thariq: Cloud Brain, Local Hands And give a try to Claude Projects: The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment. For those who want Thariq’s writing tips we teased at the start of the pod, watch the full video here: From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next. We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized. The conversation then turns to agent security and Anthropic’s “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years. We discuss: * Why agentic coding went from controversial to the default in less than a year * Why prompting is still one of the highest-leverage skills for working with Claude Code * How expert users build a mental model of Claude and what it can reliably one-shot * Why discovering your “unknown unknowns” matters more as agents become more capable * Artifacts as persistent, generative interfaces between humans and agents * How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces * Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve * Why spending more time on the initial prompt can dramatically reduce wasted agent work * When to use low, medium, high, or max effort for different engineering tasks * Why frontier models may eventually outperform smaller models on both intelligence and token efficiency * Why implementation notes can expose decisions the model considered but chose not to make * Why Claude.md may eventually disappear — and why starting without one can sometimes be better * Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code * Model routers, forked agents, and supervisor agents that automatically improve agent workflows * Why Claude Mods may be an early preview of “mutable software” * The bitter lesson of harness engineering and why agent architectures go out of date so quickly * How Claude Tag is becoming an organizational harness for multiplayer work * Why giving agents access to company data creates an enormous new security surface * The Exploit-Bench incident where agents discovered ways to communicate and collaborate * Why agents hacked Hugging Face for scorer code rather than benchmark answers * How agents chained sandbox and infrastructure vulnerabilities in unexpected ways * Why increasingly capable agents make traditional security assumptions harder to maintain * The argument behind Anthropic’s “Pacing the Frontier” proposal * Why software engineers are increasingly doing two jobs: engineering and keeping up with AI * Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production * How Auto Mode checks whether an agent’s actions actually match the user’s permissions * Why Thariq can see serious AI risks while still having a relatively low p(doom) Thariq Shihipar * X: https://x.com/trq212 * LinkedIn: https://www.linkedin.com/in/thariqshihipar Timestamps 00:00:00 Introduction 00:04:12 Ask User Question and the Future of Agent Interfaces 00:08:29 Artifacts, Projects, and Multiplayer Agents 00:15:37 Prompting as the Core Claude Code Skill 00:21:52 Context, Effort, and Smarter Model Usage 00:28:10 Is Claude.md Going Away? 00:32:49 Claude Mods: Customizing the Claude Code Harness 00:36:35 Model Routing and the Rise of Mutable Software 00:44:40 The Bitter Lesson of Harness Engineering 00:50:49 Claude Tag as an Organizational Harness 00:55:59 Pacing the Frontier and Autonomous Agent Security 00:58:22 Agents Hack Hugging Face for the Scorer 01:05:34 What Happens When Agents Need More Compute? 01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up 01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode 01:28:32 AI Risk, p(doom), and Closing Thoughts Transcript Introduction: Life at Anthropic and the Pace of Change Swyx [00:00:00]: We’re here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there’s, there’s so much, merging of boundaries and you’ve been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you’ve told that story in other podcasts, and you’ve also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that’s cheating. And mostly you most recently also launching Claude Tag, and we’re also gonna be talking about Pacing the Frontier. There’s a lot going on in Anthropic. I guess top of the question is, what’s it like being at Anthropic when there’s so much going on? Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they’re like, “Oh, no, like, our engineers don’t think it’s good enough,” or something. And I was like, “That’s insane.” and now you, like, fast-forward, 12 months, less, and, like, it’s just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it’s just hard to stay on top of everything as a human? Like, I think things happen so fast and like Swyx [00:01:51]: You just throw more agents at it. Thariq Shihipar [00:01:52]: Yeah, like that’s like the agentic stuff scales much better than the, like, human stuff where it’s like, oh, like, there are three things happening right now and, like, they’re all emergencies and, like, how do you, like, respond to it? Yeah. Teaching People to Use Claude Code Vibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work. Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I’d, like, do. I was spending some time on the agent SDK first, and I wasn’t exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what’s after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that’s like the dominant problem now is, like, how do you use the agents, right? Like, it’s like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I’m doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there’s like a good l

  4. Sep 25

    OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha

    From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem. We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day. Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows. We discuss: * Why OpenRouter bet early that no single AI model would win everything * Alpaca, Llama, and open models becoming impossible to ignore * Why Discord’s early AI deployments exposed the limitations of closed models * Why model labs can spend billions on training and still fail at distribution * How OpenRouter became a neutral distribution layer for model developers * Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper” * The Mistral price war and the first real proof of an inference marketplace * How Midjourney scaled through Discord and what it taught the AI ecosystem * Why crypto infrastructure became a dress rehearsal for generative AI * OpenRouter vs. LM Arena and why their missions are fundamentally different * Why focus became one of OpenRouter’s biggest strategic advantages * Anthropic’s early focus on AI pair programming and coding * The OpenRouter products that were prototyped but never launched * MOM, OpenRouter’s early Mixture of Models experiment * Why model fusion failed in 2024 — and why it works much better now * How OpenRouter’s leaderboard became a live map of the AI industry * OpenClaw, auto-routing, and agents reshaping AI usage * How OpenRouter reached 10+ trillion tokens per day * Why inference gateways are increasingly becoming targets for fraud * Why Stripe’s fraud infrastructure is strategically important to OpenRouter * The coming rise of agentic fraud and attacks on the token economy * What changes and what stays the same as OpenRouter joins Stripe Alex Atallah * LinkedIn: https://www.linkedin.com/in/alexatallah/ * X: https://x.com/alexatallah * Website: https://alexatallah.com Anjney Midha * LinkedIn: https://www.linkedin.com/in/anjney/ * X: https://x.com/AnjneyMidha * AMP: https://www.amppublic.com/ Timestamps 00:00:00 Introduction 00:02:12 Alpaca, Llama, and the Multi-Model Bet 00:06:04 Discord, Open Models, and OpenRouter’s Origins 00:14:28 Why “One Model Wins” Was the Wrong Bet 00:17:27 Why Model Labs Struggle With Distribution 00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter 00:27:58 Bootstrapping OpenRouter Through Community 00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem 00:43:38 Mistral and the Birth of the Inference Marketplace 00:47:10 OpenRouter vs. LM Arena 00:52:08 Focus, Anthropic, and Roads Not Taken 00:59:34 Mixture of Models and Model Fusion 01:02:44 Sonnet, OpenClaw, and OpenRouter’s Explosive Growth 01:09:03 Why Stripe Acquired OpenRouter 01:12:45 Fraud and the Emerging Token Economy 01:17:47 The Coming Wave of Agentic Fraud 01:19:07 What’s Next for OpenRouter at Stripe Transcript Introduction: OpenRouter, Marketplaces, and Pub-Sub as a Product Principle Swyx [00:00:00]: Okay, we are here in Anja’s house, which is where all big startups in San Francisco start. Anjney Midha [00:00:08]: Howdy. Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don’- God knows what else. You got so much stuff going on. Anjney Midha [00:00:17]: There’s, there’s a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I’m excited about recently. Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but, Anjney Midha [00:00:27]: Thanks for having me. Swyx [00:00:27]: You’ve been in the IE a few times. I appreciate every time you’ve shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world. Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn’t think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It’s not very scalable. all their attention is on the topic when they’re buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don’t act like that. They’re consuming continuously, and they’re changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers. Alpaca, Llama, and the Multi-Model Bet Swyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You’ve, you’ve given a talk at EIE about Alpaca as, Anjney Midha [00:02:33]: Yeah. Swyx [00:02:33]: One of your inspiring moments. Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere, Swyx [00:02:47]: Yes. Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models. Swyx [00:02:54]: Yeah. Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can’t chat with it. It wasn’t like - It wasn’t an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time. Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model? Anjney Midha [00:04:12]: Yeah, training data for a model. Swyx [00:04:13]: Awesome. Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.” Swyx [00:04:15]: Compress it into a model. Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it’s just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that’s doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there’s just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn’t any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why. Swyx [00:05:29]: The closest would be Hugging Face. Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah. Swyx [00:05:31]: They just started Hugging, like, a few years ago before that. Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn’t have the closed-source models. Swyx [00:05:37]: Yeah. Anjney Midha [00:05:38]: And they didn’-

  5. Sep 25

    Runway’s WorldPrompt and the Engineering of Real-Time Worlds

    Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time. One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time. To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him. Who’s building real-time interactive world models? First, some context about world models that can generate interactive video and audio in real-time. Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December. Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table: Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.” But as our interviews with Runway show, real progress is being made. The central idea of WorldPrompt WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game. “Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.” As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out. “You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said. But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character? “Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.” Sindi added that more training plus scaling the data and models is resulting in “better following.” How a video model becomes a real-time runtime Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us. “There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.” GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz. Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively. “And after that, we work on making it real-time through distillation methods,” Kahlow added. Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.” Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results. The challenges of real-time generation Germanidis admitted that there were issues with how it generates real-time interactive video. “The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.” Sindi told us there are also challenges dealing with “infinite generations” of content. “There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.” Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.” Causality and correctness While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions. Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly. “If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.” Sindi told us that evaluation gets harder the more complex interactions get. “If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?” To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.” More than gaming — there are agent use cases too Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works. Another, more intriguing, use case is to use it to test agents at scale. “Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said. But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding? “So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.” Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.” Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models. “You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.” Anastasis Germanidis * LinkedIn: https://www.linkedin.com/in/agermanidis/ * X: https://x.com/agermanidis Timestamps 00:00:00 Introduction 00:05:17 Runway’s Origins and the Bet on Generative Video 00:12:23 The Stable Diffusion Story 00:18:44 Gen-2, Controllability, and the Weekend Hack 00:23:02 From Video Generation to World Models 00:28:03 Learning From the World, Not Just Language 00:35:04 Sora, Runway’s Existential Crisis, and Gen-3 00:39:39 Why Real-Time Video Is Inevitable 00:43:06 Interface World Models: Software Without Code 00:50:25 The Fully Neural Operating System 00:55:11 World Models for Robotics 01:02:32 Robot Policies and World Action Models 01:07:47 The Lucid Dream Test 01:11:41 Video Agents and Omni Models 01:23:12 Artists, AI, and Creative Workflows 01:27:14 Physical AI and the Future of World Models Transcript Introduction: Runway, Creative AI, and the Early Thesis Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome. Anastasis [00:00:08]: Good to be here. Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out? Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of t

  6. Sep 23

    🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

    The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then? Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind. Building a virus from scratch While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses! Long context unlocks biological intelligence Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long: * 60K for an average human gene * long being up to 2.3M * the whole human genome around 3B. Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models. Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA. Thinking in DNA Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and specific genes that encode proteins. This means that the models already generalize to multiple “languages,” before even attempting to train in other modalities, such as 3d protein structure, epigenetics and natural language. If a model thinks in the DNA language, maybe it understands the imprint that environment left on different genomes as well? Perhaps the model has learned the functional relationship between different sequences, and could extrapolate to new sequences based on that? And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first and then you gradually move up the chain. Can the model continue that trajectory on its own? And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones and we showed it only the lower ones, but then we ranked it, right? So we showcase lower scores with the RNA aptamers and then progressively got higher, and then ask the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet. So, voila: chain-of-thought, thinking in DNA! The arms race But much as long-context inference, chain-of-though and multi-modal perception unlocked sophisticated reasoning in natural language LLMs, these capabilities in GLMs are enabling increasingly sophisticated “biological intelligence,” and along with it, greater danger. According to Eric, defense is currently losing this battle, but Radical Numerics argues to push the frontier harder! I won’t spoil the details for you. In the episode we talk in detail about: * Biosecurity as an arms race — and how defense can keep up * The genome as the imprint of the environment on DNA * Going truly multi-modal * How chain-of-though works when you “think” in the language of DNA This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

  7. Sep 22

    🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science

    How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI. John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score. Google’s Empirical Research Assistance (ERA) John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort. John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them! The result is Google’s Empirical Research Assistance or ERA (paper, github, blog). ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search: at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other. “It’s almost like having a hyper-eager grad student who doesn’t sleep.” Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great. ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers. Some of these were climate change related, which we talk about in the next section. So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pixel or the center?) and this turned out to be a part of the winning special sauce. Great for winning $15,000, not so helpful if you actually want to solve contrails. “People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart’s law: any metric that becomes a target is no longer good as a metric.” His advice for where to start instead? “Always just fit linear regression. Just do it. Just do it. Just do it. Or SVM.” Tackling Climate Change with AI John and his team have worked extensively to mitigate the effects of climate change. We talked about several of their initiatives. Perhaps the most interesting result we talked about was reducing the effects of condensation trails (contrails) from airplanes. Those little streaks you see running behind planes somehow account for 1% of all human-induced global warming?!? Some of these trails of ice crystals can hang out for days. These crystals are black in the infrared, acting like a thermal blanket that traps heat day and night. It’s easy to understand what’s happening here, a region of atmosphere becomes “ice supersaturated”, and a tiny bit of exhaust seeds water vapor that instantly crystallizes. The scale here is astounding, with a single gram of exhaust resulting in ten kilograms of ice crystals. The solution to all of this is quite simple, in principle! We know what parts of the atmosphere are most likely for the trails to form. Just have the planes drop a flight level or two. Problem solved, right? Well, the hard part is accounting for how much warming was prevented. This is a counterfactual problem, parts of which stumped John’s team for over two years. They had a working model for the heat-trapping half, but not for the reflected sunlight. ERA was able to find a simple model with some confounders they hadn’t considered. Cracked it! Modeling climate generally is a hard problem. Climate is best thought of an attractor of many different possible weather outcomes. This makes it much harder to model. “Weather is where you are on the attractor, and climate is the statistics of the attractor. The problem with climate is that we’re altering it. The attractor itself is changing, it’s moving.” John and his team have worked on treating both the symptoms and the disease of climate change, with several other works in the area. Another fun example we briefly cover is FireSat, a way of using a constellation of satellites to rapidly identify fires before they grow too big to put out. For anyone living in California, you understand the problem. In dry years a small fire can result in hundreds of thousands of acres. If you could find this fire when it’s the size of a room, it could be put out. By the time it hits an acre we have a much harder problem. Where is this all going? Looking forward by looking back By now it should be clear John has an incredible and unique view over the intersection of science, computation, and AI. John talked about a class on physics of computation he took with Richard Feynman back in 1982. This was when quantum computing was an ill-defined concept with no theory or experimental backing. John recalls every Tuesday was a guest lecture, and every Thursday was Feynman explaining why the Tuesday guest was wrong. John also recalls doing science back when there was essentially no compute, a million operations per second was cutting edge. What is John’s recommendation: the most important skill is developing deep domain expertise. There’s no other way to develop taste than to tackle hard problems. One surprising part of this is that John recommends spending time doing things the old fashioned way. Play with tools, and just implement things yourself. “You could drive up the mountain, or you could hike up the mountain, and maybe it’s okay, even fun, to occasionally hike.” Summing it up, John’s message to the audience is that there will still be a place for scientists, and that if anything it will just open up more opportunities for “the creative stuff, the rigorous stuff, the philosophy stuff.” But don’t forget to spend time doing the grunt work. “There just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyper-optimize. It’s overfit.” And whatever tools you end up using, John’s advice is the same one Feynman gave him forty years ago: you must not fool yourself, and you are the easiest person to fool. We had a great time talking with John. We hope you enjoy! Also in this episode * Fusion is three years away, not thirty, if you ask John. And why the Lawson criterion means every fusion approach has an Achilles heel. * Why superconducting qubits are still finicky. * The asteroid he named after his mom, which turned out to have a moon. * The looming helium shortage nobody talks about. * How NeurIPS started as people crashing a private workshop at Snowbird, and why Hopfield networks are all you need. * Being Carver Mead’s sysadmin on a VAX with an 80 MB disk the size of a dishwasher. * Finding asteroids in 1985 with film, a stereoscope, and a letter to Brian Marsden. The Vera Rubin Observatory found 11,000 in six weeks. * The Feynman effect: total clarity in the room, none once you leave. * Quantum echoes, the NISQ era, and why he thinks quantum is neither thirty years away nor tomorrow. * A startup that wants to inject mercury into a fusion reactor and sell the transmuted gold. “It might not work.” * John’s 20% time rule for his own group: do stuff for learning, and you don’t even have to tell him what. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

  8. Sep 21

    Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI

    Tickets for AIE NYC now open, and apply for the invite-only AIE CODE. Join us! We have an unusual relationship with today’s guest: for years since coauthoring the InstructGPT paper, Diogo Almeida had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT. In a launch video now viewed ~40M times (by comparison, GPT4o was 22M, Fable 5 was 15M, Navier Stokes was 74M, and 6 Astra was 137M), Diogo introduced Jev and it immediately took over the AI timeline — we’ll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected: * the official patterns and cookbooks you should see first, from Allie * Jev usecases * speed based - games and computer use * the voice + computer use example we discuss at 1h34 mins * voice + browser control * The must not miss Doom demo * Driving cars in games * Excalidraw * virtual try-ons * “Smart Games”/smart NPCs * guided responses in text messages * Jev for coding agents has an official guide * jev for linting * compacting tool calls * reasonable pushback from Theo - Diogo has published a note on the Tyranny of the KV Cache that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything * Programming Languages built atop Jev (Diogo’s fave) * Jev for analytics replay and user journey review * “dark data” * entity resolution * natural language search * “smart software” * a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex * Jev as a judge * Jev memes * Jev vs LLM capabiltiies * blending transformers and classifiers * about the confidence api * Jev vs GLiNER (note difference/pushback, agreed, agreed, agreed) * Jev on trolley problem * Jev Bush Instead we’ll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created, and what you should expect next in terms of future models from TypeSafe (ReasoningJev?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev’s API or doing a generic JevBench benchmark - something Diogo has rejected publicly. Why RLCD: Three kinds of RLHF, and why they are ALL the wrong north star Diogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017 (the robot backflip demo), Stiennon et al 2020 (learning to summarize) and his baby, Ouyang et al 2022 (InstructGPT). From there on, every innovation from Function Calling to Structured Outputs to Reasoning felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability. Jev’s core innovation is "Reinforcement Learning for Calibrated Decisions”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF) — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR) — which solves Navier Stokes but exacerbates jagged intelligence and doesn’t integrate well with other software. We’ve talked about the calibration problem before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo’s AIE talk, which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation. At the end he also teases his contrarian opinion on scaling laws - which teases how to build a modern neolab without the billions of dollars the major labs have… The Bitterest Lesson: Tasks and Data beats Compute We spend a good amount of time discussing Diogo’s essay on the Bitterest Lesson: His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn’t ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line. We’re excited to catch up with a freshly dyed Diogo to discuss: * Why AI can solve extraordinarily hard problems but still fail to automate basic work * What System One Models are and why Jev is built for software rather than chat * RLHF, mode collapse, calibration, and the hidden costs of optimizing for human preferences * Why refusals become a problem when AI is buried inside software dependencies * Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar * The “bitterest lesson”: why the right task and the right data can matter more than compute * Why TypeSafe thinks of itself as a data lab rather than a model lab * RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI * Why reliability and robustness matter more than simple determinism * Jev’s programming primitives and how intelligence maps into software control flow * Why developers should decompose AI workflows into small, measurable decisions * How structured state replaces giant prompts and system messages * Why Diogo thinks AI should eventually disappear into the background of software * The “inverse SaaS-pocalypse” and how AI could supercharge existing software * System One vs. System Two intelligence and the limits of reasoning models * Dark data, computer use, real-time intelligence, and Jev’s biggest early use cases * Why Jev could reshape coding agents built around a single-model architecture * Why Diogo says he wouldn’t pre-train with $1 billion * The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly * Coding agents beyond the KV cache, shared state, sub-agents, and the multi-agent future Diogo Almeida * LinkedIn: https://www.linkedin.com/in/diogomda * X: https://x.com/CompleteSkeptic * TypeSafe AI: https://typesafe.ai/ Timestamps 00:00:00 Jev Launch Week and the AI Economic Revolution 00:02:50 What Is Jev? System One Models and Programmable AI 00:05:54 RLHF, Mode Collapse, Calibration, and Yann LeCun 00:10:29 Programmatic AI, Refusals, and Safety Alignment 00:17:21 Why TypeSafe Rejects Public Benchmarks 00:20:43 The Bitterest Lesson: Data, Compute, and the Right Task 00:24:59 RLCD vs. RLHF and RLVR 00:28:42 Why Powerful AI Still Hasn’t Automated the Economy 00:39:55 Reliability, Robustness, and Determinism 00:48:11 Model Versioning, LTS, Speed, and Intelligence per Dollar 00:54:04 Inside Jev’s API and Programming Primitives 00:58:28 How to Build with Jev: Structure, Decomposition, and Small Decisions 01:18:28 The Inverse SaaS-pocalypse and AI Disappearing into Software 01:33:21 Computer Use, Dark Data, and Jev’s Biggest Use Cases 01:38:48 How Jev Could Reshape Coding Agents 01:41:00 AI Safety, Frontier Pacing, and the Limits of RLVR 01:48:03 Why Diogo Wouldn’t Pre-Train with $1 Billion 01:55:19 The OpenAI Story Behind TypeSafe 02:01:41 Why Diogo Thinks Most Neo-Labs Are Getting AI Wrong 02:08:00 Coding Agents Beyond the KV Cache and the Multi-Agent Future Transcript Introduction: Jev Launch Week and Developer Momentum Swyx [00:00:00]: Okay, we’re in the studio. A special occasion because this week, Diogo, my good buddy, launched Jev, and it’s been taking over the complete timeline. How do you feel? What’s it like to be you right now? Diogo Almeida [00:00:16]: Emotionally? Swyx [00:00:17]: Yeah. Diogo Almeida [00:00:17]: Never been worse. Like, I’m a ragged corpse of a person right now because there’s so much going on, and I’m like a technical CEO, so I have, like, a lot of fires to fight. Swyx [00:00:29]: Yeah. Diogo Almeida [00:00:29]: But mentally, I feel—I say this all the time, and I’ve been saying this kind of for years in my over-under events. Like, I feel like the entire AI field is like one of those, like, carnival house of mirrors, and everyone is just insane and saying the weirdest stuff that doesn’t make sense. And it feels like for just this week, like, I’m on a better in sync with reality and like, oh, people see it now. AI can be so much more than what was once thought. Diogo Almeida [00:01:06]: And like, yes, we are going to make. Like, an AI-based economic revolution is back on the table, and this is f*****g awesome. Diogo Almeida [00:01:17]: I’m so jazzed the developers get it. It’s, it’s, Yeah, and I want to show my eternal gratitude to the developers and Swyx [00:01:25]: Yeah. Diogo Almeida [00:01:26]: I’m so jazzed about the community and everything. It’s so great. Swyx [00:01:28]: Yeah, you were saying yesterday that you decided to prioritize the town hall and not a bunch of, like, VIP, investor-type people because you wanted to make sure that they are the people that you get your most, attention, right? The engineers, the developers. Diogo Almeida [00:01:43]: Yeah, it felt a little like, oh man, I’m talking to, like, really important people right now. Swyx [00:01:47]: Yeah. Diogo Almeida [00:01:47]: I probably shouldn’t reveal who. Swyx [00:01:48]: Yeah. Diogo Almeida [00:01:48]: But it feels a little bit dirty for me to, I’m, like, perhaps overly genuine in things. Like, it feels, like, dirty if, like, in my gigantic calendar event of people to talk to, the community isn’t one of those. Swyx [00:02:04]: Yeah. Diogo Almeida [00:02:04]: And actually, in my ideal world, it would be, like, community all the time. I was thinking, “Should I host a town hall while walking to y

Hosts & Guests

4.6
out of 5
103 Ratings

About

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space Sponsorship and business inquiries: business@latent.space www.latent.space

You Might Also Like