Latent Space: The AI Engineer Podcast

Latent.Space

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space Sponsorship and business inquiries: business@latent.space www.latent.space

  1. 1d ago

    Simulation: the new Scaling Law — Joon Sung Park, Simile AI

    When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working! From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth. We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs. We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one. We discuss: * How Smallville and Generative Agents led to Simile * Why Joon’s team asked: “What if we can just recreate the world that we live in?” * Why useful personal agents require deep models of their users * Memory architectures, Markdown files, and the limits of prompting * “Social physics” and behavioral foundation models * Why web data captures what people say more than what they actually do * Interviews, transactions, observational data, and randomized controlled trials * Why predicting the future matters less than understanding how to shape it * How Simile creates representative simulated populations * Simulation versus prediction and the connection to Foundation’s psychohistory * How to evaluate simulations instead of simply stacking LLM hallucinations * Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy * Why frontier models can struggle to reproduce real human behavior * Why good simulations need to reproduce human biases and mistakes * Post-training models on randomized controlled trials * Population-level versus individual-level simulation * Scaling laws for human simulation * The long-term ambition to simulate all 8 billion people on Earth * Whether simulations could help solve climate change or detect collapsing democracy * Thomas Schelling and the history of agent-based modeling * Why future simulations could require an entire data center * Multi-agent simulations and what happens when simulated people interact * Replacing expensive human panels with synthetic populations * Why market research is only the starting point for simulation * Why Joon sees simulation as surprisingly similar to painting * Using simulation to study questions like UBI * Whether we are already living in a simulation * Why AGI and simulation may be the twin technologies of advanced civilizations Joon Sung Park * LinkedIn: https://www.linkedin.com/in/joonspark * X: https://x.com/joon_s_pk * Website: https://www.joonsungpark.com * Simile: https://www.simile.com Timestamps 00:00:00 Introduction and Joon’s Path from Art to AI 00:01:46 Smallville, Generative Agents, and the Origins of Simulation 00:05:03 “Let’s Just Create a World” and the Future of Personal Agents 00:09:53 Social Physics and Behavioral Foundation Models 00:14:08 Prediction vs. Simulation: How Do You Shape the Future? 00:16:59 How Simile Models Real People and Populations 00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy 00:30:23 Post-Training Models to Reproduce Human Behavior 00:40:04 Scaling Laws and Simulating 8 Billion People 00:43:10 From Schelling to Society-Scale Agent Simulations 00:46:13 The Cost and Economics of Simulating the World 00:52:05 Real-World Use Cases, Synthetic Populations, and the Market 00:57:27 The Future of Simulation, Painting, and UBI 01:04:23 Are We Already Living in a Simulation? 01:06:08 Building Simile and Hiring Transcript Introduction: Joon Sung Park, Simile, and the Story So Far Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here? Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy. Vibhu [00:00:49]: Painting. Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that’s what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am. Smallville, Generative Agents, and the 2023 Breakout Paper Swyx [00:01:46]: So there’s a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper. Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats. Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m not sure. Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast. Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations. Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times. Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you’ve read recently?” It’s this one. Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers. Foundation Models and the Search for Killer Applications Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It’s really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came together Swyx [00:03:35]: Who coined foundation models. Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn’t, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We’ve known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It’s social m

  2. Aug 11

    🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery

    This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at the heart despite being all of 2 years old. The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story! Editor’s note: not to be confused with Chai AI, which was another top pod of ours. Pharma suddenly doing big AI tools deals For the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing. Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios. The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust. Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed. Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this. RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much. Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that? Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic. This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition. Photoshop for molecules With that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts. What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with. — Neil Patil, (Chai product lead) This means better UX, such as a molecule editor that is more like a CAD or graphics design program than a chatbot. Their approach has paid off: since June, Chai has announced three more major deals: Lilly, Novartis, argenx, plus an expansion of their Eli Lily program. This episode is too full of quotable moments for a short blog, so tune in to learn about * Why protein tokens have the highest downstream value of any token * Climbing levels of abstraction as models improve * How Pharma, VC, and research are all just portfolio optimization * How better tech changes the whole portfolio * How relentless focus on simplicity leads to scale Plus much more! This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

  3. Aug 3

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    Watch the full episode on YouTube: We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3: And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF: Three years ago, inference engineering barely existed as a category. Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem. In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model. The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them. We discuss: * What happens when a 200,000-token request enters an inference system * Cache-aware routing and reusing previously computed KV cache * Why prefill and decode are increasingly handled by different GPUs * When dedicated deployments become cheaper and more reliable than shared APIs * How speculative decoding uses a smaller model to accelerate a larger one * Tool calling, structured outputs, and what LLMs actually do * What it takes to support a new open model on day zero * Grafting Kimi’s vision encoder onto GLM-5.2 * Retrofitting inefficient model layers with components from other architectures * Why models sometimes collapse into repeating the same token * How hardware, kernels, and race conditions create nondeterministic failures * Preserving model fidelity while making inference faster * How quantization errors can cancel each other out * Why inference optimizations still deliver gains of 20%, 100%, and 200% * How optimized serving can make a model up to 10× faster * NVIDIA Dynamo, KV-aware routing, and distributed model serving * Speculative decoding the speculative decoder * Why local AI is about making models less dumb while data-center AI is about making them less slow * Tensor, expert, and pipeline parallelism across GPUs * Hardware-aware model design, auto-tuning, and the case against mega kernels * Rubin and why inference is becoming a systems problem * Whether modern GPUs are evolving into programmable AI ASICs * Why enormous models like Kimi K3 require GB300-class hardware * Why open-source video generation still trails Veo, Kling, and other closed models * The quadratic attention bottleneck behind long-form AI video * Autoregressive video, real-time generation, and compounding quality drift * Why future video systems may combine autoregressive and diffusion architectures * Training for inference and inference for training * Continuous post-training, deployment, evaluation, and improvement loops * How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself * Why faster networking could unlock dramatically faster decoding * Continual learning, KV-cache compaction, and persistent model memory Show Notes * How to build a day-0 API for Kimi K3 * 22580: From GPT2 to Kimi3, Explained Philip Kiely * LinkedIn: https://www.linkedin.com/in/philipkiely * X: https://x.com/philipkiely * Inference Engineering: https://www.baseten.co/inference-engineering/ Ali Taha * LinkedIn: https://www.linkedin.com/in/aliestaha/ * X: https://x.com/waterloointern Timestamps 00:00:00 Introduction and the 200K-Token Prompt 00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling 00:11:26 Launching Production-Ready Open Models 00:19:06 Model Retrofits, Failure Modes, and Nondeterminism 00:28:22 Quantization and Canceling Errors 00:32:15 The Race to 10× Faster Inference 00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI 00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels 01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips 01:10:03 Giant Models and the Limits of GPU Memory 01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation 01:21:47 Audio, Images, and Diffusion Models 01:27:32 Training, Self-Optimizing Models, and Continual Learning 01:40:06 Closing Thoughts Transcript Introduction: Baseten, Waterloo Intern, and Inference Engineering Swyx [00:00:00]: Okay, we’re here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you’ve done, you and I have done before, as well as Ali. Welcome. Ali [00:00:15]: Pleasure to meet you. Swyx [00:00:15]: Waterloo intern. Ali [00:00:16]: Waterloo intern, always. Swyx [00:00:17]: When did you get “Waterloo intern” as a handle? Ali [00:00:19]: As a handle? Oh. Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.” Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer. Philip [00:00:30]: So we have to figure out who’s gonna get the handle. Ali [00:00:33]: Well, I’ll pass the torch over to the next intern. Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad. Ali [00:00:37]: To another Waterloo intern. No, bruh. Philip [00:00:39]: Yeah. Ali [00:00:39]: Intern. Swyx [00:00:40]: Intern, yeah. Ali [00:00:40]: And no. Philip [00:00:41]: You gotta get an intern from Waterloo. Ali [00:00:42]: Yeah, I’ve gotta get an intern from Waterloo. Swyx [00:00:44]: Right. Ali [00:00:44]: But they have to follow the path. Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it’s like whoever Baseten gets from Waterloo. Ali [00:00:48]: Right. Swyx [00:00:49]: Has the title of Waterloo. Ali [00:00:50]: It stays in the ecosystem. Philip [00:00:51]: Exactly. Ali [00:00:52]: Halfway through the internship, you either get it or you’re out. Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle. Ali [00:00:59]: Just say it. Philip [00:00:59]: For everybody. Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you’re an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten’s inference? What’s the process of query through GPU model routing, balancing, all that? What is all the stuff that we don’t think about? Long Context Requests, KV Cache, and Cache-Aware Routing Philip [00:01:26]: With a long query specifically, the first thing that I’m gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it’s gonna be a lot easier for me and a lot cheaper for you. So the first thing that we’re gonna look at is some cache-aware routing, where we’re going to see, we probably have a number of instances, a number of replicas up serving whatever model you’re hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you’re doing two hundred thousand tokens, it’s probably coding or a multi-turn agent or something where you would expect to have that cached. If you don’t, we’re gonna have to send it to a prefill worker. We’ve at least on certain models disaggregated prefill and decode, so you’re going to have one set of GPUs that’s solely going to process the input, create the KV cache, and get you your first token, and then that’s going to be passed over to a separate set of GPUs, which is going to run decode. We’re going to iteratively make those tokens. We’re probably going to have some speculator model in front of that. I’m going to assume that you’re doing coding, and because of that, our speculator model, which assumes you’re doing coding, is gonna have a high draft token acceptance rate. If I’m wrong and you’re asking me to summarize every Harry Potter book, it’s gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?” Swyx [00:03:04]: Except Ba

  4. Jul 28

    Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI

    There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right. A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents: We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March. With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team. However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and artifact needed to reach it. From building no-code products at Airtable to leading Productivity Engineering at OpenAI, Akshay Nathan has spent much of his career trying to make the power of software accessible to people who do not write code. In this episode, Akshay joins swyx and Vibhu to unpack the launch of ChatGPT Work, why Codex unexpectedly took off among non-developers inside OpenAI, and the company’s broader plan to bring useful agents from software engineers to knowledge workers and eventually everyone. We go deep on the shared agent harness behind Codex and ChatGPT Work, why OpenAI brought the experiences together without making them identical, and how persistent computers, artifacts, Sites, plugins, memory, and sub-agents are changing what people can delegate to AI. Akshay explains why some teams are replacing decks and spreadsheets with interactive websites, how agents can gather context across code, Slack, documents, and local files, and what OpenAI learned from personal-agent products like OpenClaw. Side note: also don’t miss Abhihek’s sandbox track keynote at AIE, which now powers a lot of the sandboxing for ChatGPT Work… and yes was also broken by an unreleased OpenAI model in the recent HuggingFace incident. Akshay also reflects on how AI is transforming product development itself: why more people will become generalists with a specialty, why ideas and taste become the bottlenecks when almost anyone can build, why LLMs still struggle to generate genuinely grounded new ideas, and why teams must distinguish increased motion from actual progress. We discuss: * Why Codex unexpectedly took off among non-developers inside OpenAI * Why employees felt like using Codex gave them a new superpower * The product insight that led OpenAI to build ChatGPT Work * Why Codex and ChatGPT Work share the same underlying agent harness * How their UX, Git visibility, artifacts, and sandboxing defaults differ * Why OpenAI merged its agent experiences instead of building separate products * How AI is blurring the boundaries between engineering, design, strategy, and operations * Why OpenAI wants the default model configuration to work for most users * When power users should use deeper reasoning, Ultra, or multi-agent modes * Artifacts, agentic spreadsheets, and creating high-fidelity work products * Why interactive Sites may replace decks and spreadsheets * The challenge of designing a simple interface for an agent that can build almost anything * Why users should retry tasks that models could not handle three or six months ago * How AI can gather context for performance reviews without replacing human judgment * The OpenAI automation that turns internal Slack and document activity into memes * What reaching ten million ChatGPT Work and Codex users means for the product * How OpenClaw inspired persistent environments, scheduled tasks, and personal agents * Using ChatGPT for financial planning, budgeting, workouts, meals, and household management * The design tradeoffs behind sub-agents and how much of their work users should see * ChatGPT memory, Chronicle, and long-term context * Why AI may make more people generalists with deep specialties * Why ideas and taste become more important when almost anyone can build * Why LLMs still struggle with the instruction “bring me new ideas” * Measuring productivity through quality at-bats instead of commits, tokens, or pull requests * The critical difference between AI-generated motion and meaningful progress Akshay Nathan * LinkedIn: https://www.linkedin.com/in/akshaynathan/ * X: https://x.com/akshaynathan_ Timestamps 00:00:00 Introduction and Bringing the Power of Code to Everyone 00:01:33 Joining OpenAI and Preserving a Startup Culture 00:02:40 What OpenAI Learned from Enterprise AI Adoption 00:05:28 Why OpenAI Built ChatGPT Work 00:07:17 Codex vs. ChatGPT Work and the Shared Agent Harness 00:12:07 Why OpenAI Merged Its Agent Experiences 00:16:24 Models, Reasoning Levels, and Choosing the Right Default 00:20:26 Artifacts, Agentic Spreadsheets, and Model–Product Collaboration 00:24:22 Why Sites Could Replace Decks and Spreadsheets 00:30:08 Designing an Agent That Can Build Almost Anything 00:34:28 From Developer Agents to Knowledge Work—and Everyone 00:36:07 Power-User Advice and AI-Assisted Performance Reviews 00:40:41 OpenAI’s Internal AI Memes and the Ten-Million-User Launch 00:44:39 OpenClaw, Personal Agents, and ChatGPT as an Operating System 00:50:24 Sub-Agents, Ultra Mode, and How Much Control Users Need 00:54:39 ChatGPT Memory, Personalization, and Chronicle 01:00:19 How AI Is Reshaping Product Development and Tech Roles 01:03:15 Ideas, Taste, and Why LLMs Struggle to Generate New Ideas 01:04:42 Measuring Productivity, Quality At-Bats, and Motion vs. Progress Transcript Introduction: Akshay Nathan, ChatGPT Work, and the No-Code Arc Swyx [00:00:00]: We’re here in the studio with Akshay from OpenAI. Welcome. Akshay Nathan [00:00:07]: Thank you. Swyx [00:00:08]: And with our trusty co-host, Vibhu. So you recently launched ChatGPT Work. You lead Core Product Engineering. It’s been a long journey, into all this. I find it very interesting that you started with no code or low code, with Walrus and Airtable. And to some extent, ChatGPT Work is like the super app of super apps of, well, here is the ultimate no code. You just write a prompt. Akshay Nathan [00:00:32]: Yeah. It’s funny how things come, full circle. I think for a long time in my career, I started my career working consumer fintech, but then after that, like, there’s this hypothesis that, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more, accessible way, then that would be truly magical. We were working on a startup. It’s funny, like, before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kinda jank, back then, but doing what we can, and then worked at Airtable for a while on the same thesis that, like, if we can bring a database or the primitives behind a database to people, that’d be really useful to them. But once LLMs came onto the scene, it became clear that, this was the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what’s going on underneath the hood. And so, like, I think this launch and a lot of the stuff that we’ve been up to is, like, the manifestation of that. From Walrus and Airtable to OpenAI Vibhu [00:01:33]: How was stuff when you joined? So you joined OpenAI 2023. Now we’ve got, so much more stuff, so ChatGPT, Codex app, ChatGPT Work. Have things changed? Joining OpenAI and What Hasn’t Changed Akshay Nathan [00:01:44]: I think the more interesting thing is how things haven’t changed. Like, one, I joined I remember when I joined, it was, like, five hundred people. One thing I was worried about was, like, I was looking for something, more early stage and, like, was it gonna feel startup enough? And I joined, and I was like, “This feels even more startup-y than I could ever imagine.” And, like, that really hasn’t changed even till now. I think the, like, level of, like, bottoms-up ambition and, like, the ability of anyone to, like, do anything or have an idea and ship it is really cool. But on the, like, mission side, I think what was really compelling to me is this mission of, bringing frontier intelligence to everyone. Like, building AGI and then bringing it to everyone. And, I think acknowledging back then that, like, that vision is gonna, not be a linear progression. Like, we’re probably gonna, like, try different products and have different things that succeed and don’t. But the vision has stayed the same, and the mission has stayed the same, and we’re starting to see the pieces, fall together, and that’s really cool. Enterprise Lessons: No One-Size-Fits-All AI Swyx [00:02:40]: You worke

  5. Jul 23

    Inside the Model Factory — Eiso Kant, Poolside AI

    In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines’ recent release nearly 10 times their size. Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club: From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five. We go deep on Poolside’s Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web. We also discuss model-harness co-design, Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise, open-source AI, regulation, NVIDIA and TSMC’s influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside. We discuss: * How Andrej Karpathy’s RNN work inspired Eiso to start building language models for code in 2015 * Why Eiso spent four years and $12 million pursuing an idea before the market cared * Why ChatGPT felt like vindication and brought Poolside back to open source * Why Eiso would prefer 100 foundation model companies over an oligopoly of five * The difference between releasing open weights and publishing genuinely open research * Why Poolside deliberately built a global research organization outside the Bay Area talent war * Why model building is ultimately 90% engineering * The Model Factory: Poolside’s end-to-end system for rapidly training and improving models * How fewer than 70 researchers run roughly 10,000–20,000 experiments each month * How Poolside moved from six-month model cycles to five- and eight-week launches * Why streaming data directly into training unlocked faster experimentation * How immutable data, versioned code, and reproducibility enable rigorous model research * Why Eiso wants capable researchers to leave their labs and become Poolside’s competitors * Why 95% of model building can be reduced to better data or compute efficiency * Laguna S and why persistence, verification, and backtracking can outperform raw intelligence * Why smaller models may handle far more knowledge work than previously expected * Why reinforcement learning will move earlier into pre-training * Why next-token prediction is still failing to extract enough knowledge from the web * Why distillation and environments have become the AI industry’s favorite “drugs” * Why mid-training is really an early form of curriculum design * Low-precision training, networking bottlenecks, and the next gains in compute efficiency * Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch * Why model builders can often evaluate a new checkpoint within its first 30 minutes * Model versus harness: where agent capabilities actually come from * Why Poolside sees coding and long-horizon software tasks as a path to AGI * Why Eiso thinks MCP and traditional tool calls are “stupid” * Why future agents will write scripts instead of choosing from dozens of predefined tools * The case for minimal harnesses, containers, and model freedom * Why Poolside is prioritizing vision but does not expect to work on audio soon * Why language may be the most compute-efficient modality for encoding knowledge and reasoning * The real cost of model development and why the final training run is anticlimactic * The story behind the Poolside name and why it represents refusing to lower ambitions * How Poolside raised $500 million while investors still questioned whether AGI was real * Why intelligence could become the world’s most demanded and commoditized resource * When open models may become too capable to release without restrictions * Why unilateral AI safety does not work in a globally competitive environment * How regulation could accidentally lock in an oligopoly of two or three AI companies * NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress * Why reinforcement-learning wall-clock time is one of Poolside’s biggest bottlenecks * Why Poolside trains models from scratch instead of simply distilling larger models * How AI changes the way companies should measure engineering productivity * Why agency may become the most important quality for employees in the AI era * How leaders align high-agency people through shared goals and clear constraints * Hiring across research, post-training, pre-training, architecture, evals, and engineering at Poolside Eiso Kant LinkedIn: https://www.linkedin.com/in/eisokant X: https://x.com/eisokant Poolside: https://poolside.ai Timestamps 00:00:00 Introduction 00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers 00:02:26 The $12M Failure and ChatGPT Vindication 00:03:39 Open Source and the Case for 100 Foundation Model Companies 00:09:22 Open Weights, Open Research, and Poolside’s Global Team 00:16:04 The Model Factory: Why Model Building Is 90% Engineering 00:20:19 Agents, Automated Experiments, and Early Signs of RSI 00:24:04 Streaming Data, Reproducibility, and Scientific Rigor 00:30:35 Creating More Foundation Model Companies 00:36:07 Laguna S: Persistence vs. Raw Intelligence 00:43:01 Reinventing Pre-Training, RL, and Curriculum Design 00:52:33 Low-Precision Training and Squeezing More From Smaller Models 00:58:37 Model Harnesses, Coding Agents, and the Path to AGI 01:09:26 Why MCP and Traditional Tool Calls Are “Stupid” 01:13:04 Vision, Multimodality, and Why Language Still Matters 01:18:15 Scaling Models and the Real Economics of Training 01:20:40 Why Poolside Is Called Poolside and Raising $500M 01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly 01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck 01:41:52 Smaller Models, Distillation, Engineering Productivity, and Hiring Transcript Introduction: Eiso Kant, Poolside, and Open Models Swyx [00:00:00]: All right, we’re here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome. Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here. Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I’m on my way to SF.” I was like, “You’re on a plane right now, right?” Like, hey. Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let’s do it. Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don’t have to prepare that much because if you’re truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don’t live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we’re gonna talk about. But I guess, like, what got you into democratization of AI? Like, it’s not obvious from your LinkedIn or something. From Karpathy’s RNN Post to Sourced Eiso Kant [00:00:57]: No, it’s not at all. I don’t think it’s obvious how I got in this space. I owe getting into this space to Andrej Karpathy. Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.” Swyx [00:01:10]: Neural Nets, yep. Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There’s a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn’t. and there’s a little-- There’s an example of code a little bit further down. Yeah, so Shakespeare. Swyx [00:01:47]: Shakespeare. Swyx [00:01:49]: Cool Eiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool

  6. Jul 21

    🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

    Bet on information If test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data. Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need more information rich data. This is what Chu and Bo’s teams have done, and here is what ~30x the information buys you: Now we can scale with parameters and training compute! We don’t know how much this effort costed, but we can guess that data collection experiments and infrastructure was a few tens of millions, and compute + headcount + research was a few million. The budget looks like a RL rollout budget, rather than a data rich pre-training one. We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting that information rich data is the key to AI-driven drug development. Chu was recently promoted to Chief Discovery Officer and Bo to Chief AI Scientist, underscoring just how strategic Xaira considers this bet. Reverse engineering the human cell If you had to figure out how a human cell works, what would you do? A good place to start might be by documenting what genes are expressed (e.g. what RNA is floating around) in different kinds of cells, in different circumstances. That is CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute that maps each cell to a count of how many times 20K-30K genes were detected in that cell, plus detailed metadata about every cell. A ~4 trillion-entry matrix. If the Protein Data Bank (PDB) unlocked structural biology models (Boltz Episode, ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virtual Cell models. Bo Wang built one of the most influential, scGPT, that became the starting point for Xaira’s new model. RNA expression ≠ Virtual Cell Models trained on CELLxGENE describe the relationship between cell types and cell states, but they are not good at predicting what will happen if we make changes to RNA expression. Changes in gene expression are highly correlated, and its is difficult (impossible) to figure out what causes what in most cases. If you could “turn the dial down” on one gene at a time, however, then you would be able to observe what is upstream and downstream of a given gene. You could tell if A → B & C or B → A & C or B → A, C → B → … If you did this for all of the genes, then maybe you could train a model that could predict what would happen to a cell if you change a gene (e.g. with a drug or a gene edit). Or maybe you could figure out the least invasive way to change a particular gene’s expression. X-Atlas → X-Cell This is exactly what Chu and Bo’s teams have done. The data set is called X-Atlas and the model is called X-Cell. In this episode, we discuss: * Why the team abandoned autoregression for diffusion * The CRISPR-based experiments that run millions of tests in parallel, and generate the raw data for X-Atlas and X-cell * Generalization to real lab experiments in real human cells * Beating the linear baseline that has outperformed previous models * Justifying a kitchen-sink of priors, and how that stacks up vs. data and architecture Bo also shared with us some of the (major) advantages he has as an academic vs. industry leader, and how his labs keep up with the breakneck pace of AI innovation. Check out the full episode on YouTube, or your favorite podcasting platform! This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

  7. Jul 16

    🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences

    Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward a scientific superintelligence. Their automated lab is almost hypnotizing to watch. They have floating plates zipping around on Wall-E-esque tracks, used vision-language models to control Windows 95 boxes, and created the world’s largest collection of voided warranties. In the process they’ve built a massive library of scientific reasoning tokens. Over 10 trillion of them, all experimentally validated. No warranties were voided in the making of this video To say Lila is ambitious is an understatement. Their goal is a scientific superintelligence wired directly into the wet lab. They are all in on the bitter lesson, and the thesis follows from it: a lab is an infinite token generator. Produce data at scale, and the synergies give you a general reasoner that can tackle any scientific problem. They are committing hard. Biology, chemistry, drug discovery, and materials science, all at the same time. Time will tell if it works, but it is an exciting hypothesis. In our latest episode we sat down with Lila’s very own Andy Beam (CTO) and Rafa Gómez-Bombarelli (CSO, physical sciences) and went on a journey through the possibilities of AI-run science, almost as wide-ranging as Lila’s goals. Did we mention they do both materials science and biology? In the same AI science factory? Same time, same lab, same AI. Finally a guest who can settle a long-running debate we’ve had amongst ourselves: is biology or materials science harder? Watch to find out! We discuss: * The internet is spent, science is next. Why Lila thinks the scientific method is the last untapped internet-scale dataset, and why they treat RL as a data generation mechanism with nature as the verifier. * The lab as a data center. Instruments as nodes on a graph, a magnetically levitating “PCI bus” transport layer between them, orchestration as a slurm queue. Andy is not short on analogies. * Why Lila insists it is not an automation company. They optimize for flexibility and generalizability over raw throughput, which means humans stay below the API line wherever automating does not pay. * Your experiment has a runtime. We put Escalante Bio’s question to Andy: if science is the token generator, what is the runtime of your data collection? His answer, in short, is that you cannot make the ribosome go faster. Why Lila bets on fast round-over-round iteration rather than big noisy multiplexed screens, and how Rafa’s team rebuilt a gas sorption measurement to run roughly 2,500x faster. * What is actually in 10 trillion scientific tokens. Not sequences. Experimentally verified reasoning traces, a kind of data that Andy argues exists on the internet in quantities that round to zero. * Breadth as a path to depth. Small molecule chemistry priors transferring to metal organic frameworks for carbon capture, and the claim that the general model beats domain-specific models sample for sample. * If you have the data, what do you need the model for? Sri Kosuri’s koan about the ML-for-drug-discovery business model, and Andy’s answer: the coding model got better because it also read Shakespeare and carnitas recipes. * The serendipity they want to automate. Emily Whitehead survived the first pediatric CAR-T cure only because the doctor treating her happened to know, from pediatric arthritis, which antibody would blunt her IL-6 response. Roll that dice again and you probably lose her. Breadth is how you stop depending on luck. * Move 37 for catalysts. Model suggestions for platinum-group-free electrocatalysts that went from boring, to what a 40-paper expert called stupid, to the best performers they have made. * Six months to in vivo CAR-T data in non-human primates, and the zero-FTE virtual startup commercial model that fell out of it. For context on why that number is startling, AbbVie paid $2.1B for Capstan on the strength of preclinical in vivo CAR-T data. * You cannot have scientific superintelligence if you are just a good test taker. Ken Stanley, who wrote Why Greatness Cannot Be Planned, runs open-endedness at Lila. RL at scale gives you a ruthlessly Vulcan problem solver. Machine creativity is a different thing, and it is the part nobody has solved. * The chain of thought is an unreliable narrator. The model reasons in latent space and only emits tokens. Sometimes it skips the experiment entirely and is still right. So how much do you trust the reasoning versus the verifier? * Reward hacking when the rollout is physical. Chains of thought that collapse into repetition, and a model that got annoyed and swore at the scientist who kept asking it to redo a plate map. What happens when a pathological loop has a wet lab inside it? * The bittersweet lesson. Rafa’s inversion of the bitter lesson: in AI, scaling is a roadmap. In materials, scaling is a filter, because only the things that scale end up mattering. * Not your typical Flagship company. Why a famously single-asset biotech incubator spun out a platform bet, and Andy’s line that if Lila called itself a biopharma it would have a top-three GPU cluster. * Bottlenecks they would remove by fiat. Sim-to-real for physics-based simulation, and the fact that RL training runs at roughly 5% mean FLOP utilization. Watch on YouTube: This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

    🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
  8. Jul 8

    Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

    We’ve been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we’re excited to conclude this series returning to Modal, which has just raised a monster $355M Series C. The cloud was built for developers. But agents are now changing that. The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads. However, agents don’t have that luxury. Now in this new era of agents, everything has to be tighter. They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research: Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time: At the time, Modal was just a teeny little company with a $17M Series A. Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate. In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don’t fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience. We go deep on Modal’s AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal’s capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again. We discuss: * Why Kubernetes wasn’t built for bursty AI workloads * How Modal started as a better runtime before becoming an AI cloud * Why Modal added GPUs before ChatGPT * The shift from developer experience to agent experience * Why observability matters when agents are writing the code * Elastic inference for custom models across audio, video, robotics, and comp bio * GPU snapshotting, cold starts, and why inference workloads are so bursty * Why RL rollouts can require 100,000 sandboxes * DeFlash, speculative decoding, and frontier-level inference performance * Auto Endpoints and making optimized inference easier to deploy * What Modal adds beyond vLLM, SGLang, and raw GPU rental * Modal’s 17-cloud capacity pool and supercloud strategy * Networked sandboxes, sidecars, private IPv6, and RDMA * Serverless multi-node training for post-training and research workloads * Auto-research, model-guided sweeps, and agents launching GPU experiments * Compute strategy, capacity planning, and batch tiers * Why production agents need specialized sandboxes and hard guardrails * Modal’s take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal Bench Akshat Bubna * LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103 * X: https://x.com/akshat_b Modal * Website: https://modal.com Timestamps 00:00:00 Introduction 00:00:39 Modal’s origin and why Kubernetes wasn’t enough 00:04:32 Developer Experience → Agent Experience 00:06:21 Modal’s AI cloud primitives 00:09:14 Sandboxes, agent loops, and proto-Cognition 00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes 00:15:24 DeFlash, speculative decoding, and Auto Endpoints 00:19:59 Production-grade inference beyond raw GPUs 00:22:00 Background agents, Ramp Inspect, and the agent lifecycle 00:24:08 Modal’s 17-cloud supercloud strategy 00:26:40 Networked sandboxes, private IPv6, and RDMA 00:32:48 Multi-node training, post-training, and auto research 00:37:36 Compute strategy, capacity planning, and batch tiers 00:40:55 Open models, real-time AI, and production agent infra 00:43:06 Hard guardrails, managed agents, and specialized sandboxes 00:46:06 Why AI made infrastructure exciting again 00:48:30 Model APIs, differentiated products, and agentic video 00:51:50 CI, coding-agent infra, SDKs, and Modal Bench 00:57:28 Closing Thoughts Transcript Introduction: Modal, Series C, and the Art Party Swyx [00:00:00]: We’re here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C. Akshat [00:00:10]: Thank you. Swyx [00:00:11]: Your party yesterday was amazing. Akshat [00:00:15]: Yeah. Swyx [00:00:15]: From all the photos and all the swag. Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin. Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story. Modal’s Origin: A New Runtime Beyond Kubernetes Akshat [00:00:39]: I first met Eric, who’s the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It’s because you have to run them on Kubernetes. Kubernetes is hard to manage. It’s not built for burstiness and, custom images, Swyx [00:01:03]: Yeah Akshat [00:01:03]: It has a terrible developer experience. Swyx [00:01:05]: And I’ll, I’ll interject Akshat [00:01:06]: Yeah Swyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there’s a bit more of the story there from Spotify and all those things. Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment. Akshat [00:01:26]: Yeah. Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines? Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it’s a very useful primitive in itself. It’s There’s a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would’ve been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out. From Serverless Containers to GPU Workloads Swyx [00:02:19]: Nice. Akshat [00:02:19]: We just didn’t think it would be that big of a deal. Swyx [00:02:22]: Yeah, just like add A100. Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it? Akshat [00:02:28]: Yeah. Primarily it’s just, none of the tooling that was out there was built for, one, a really great developer experience, and also there’s a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there’s just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step. Software-Defined Infrastructure and Decorator-Based DX Swyx [00:03:13]: Yeah. Yeah. Be nice. I don’t know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that. Akshat [00:03:22]: Yeah, the self-provisioning Swyx [00:03:23]: Self-provisioning. Akshat [00:03:24]: Yeah. Swyx [00:03:24]: Yeah. I can’t even remember my own post. Swyx [00:03:26]: And then you put me on the landing page. Akshat [00:03:28]: Yeah. We really like, the term and so we stole it. Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right? Akshat [00:03:37]: Yeah. Swyx [00:03:37]: Was that a big part of the original Akshat [00:03:39]: Yes Swyx [00:03:39]: Story or it was just like a DX layer? Akshat [00:03:41]: That was, really important because we really didn’t want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you’re doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that’s more expressive and dynamic. and so yeah, that was always a very important part. Swyx [00:04:04]: Then the pushback is this is a DSL. Akshat [00:04:07]: Yeah. Swyx [00:04:07]: It’s you’re closed source. I am locked into Modal. Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and su

Hosts & Guests

4.6
out of 5
102 Ratings

About

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space Sponsorship and business inquiries: business@latent.space www.latent.space