• Rechercher
  • Accueil
  • Nouveautés
  • Classements

Technologies

  • Latent Space: The AI Engineer Podcast
    Latent Space: The AI Engineer Podcast

    1

    Latent Space: The AI Engineer Podcast

    Latent.Space

  • Version History
    Version History

    2

    Version History

    The Verge

  • 9to5Mac Happy Hour
    9to5Mac Happy Hour

    3

    9to5Mac Happy Hour

    9to5Mac

  • The a16z Show
    The a16z Show

    4

    The a16z Show

    Andreessen Horowitz

  • Acquired
    Acquired

    5

    Acquired

    Ben Gilbert and David Rosenthal

  • Lenny's Podcast: Product | Career | Growth
    Lenny's Podcast: Product | Career | Growth

    6

    Lenny's Podcast: Product | Career | Growth

    Lenny Rachitsky

  • Telecoms.com Podcast
    Telecoms.com Podcast

    7

    Telecoms.com Podcast

    Telecoms.com

Les indispensables

  • برق مع عبدالله السبع
    Actualités technologiques
    Actualités technologiques

    15/10/2020

  • Lew Later
    Technologies
    Technologies

    Tous les jours

  • Tech Life
    Technologies
    Technologies

    Chaque semaine

  • Babbage from The Economist
    Sciences
    Sciences

    Chaque semaine

  • Decoder with Nilay Patel
    Affaires
    Affaires

    Chaque semaine

  • Darknet Diaries
    Technologies
    Technologies

    Tous les mois

  • TED Radio Hour
    Technologies
    Technologies

    Chaque semaine

  • The Self-Improving Company | Kavak's AI Playbook

    -23 h

    The Self-Improving Company | Kavak's AI Playbook

    Angela Strange and Gabriel Vasquez are joined by Alejandro Maza Ayala, Chief Product & AI Officer at Kavak, to unpack how the Latin American used-car marketplace rebuilt itself around AI agents, with 96% of customer interactions and 95% of transactions now handled by agents. Alejandro explains why Kavak decided that simply giving employees AI tools wasn't enough, and instead redesigned the company's systems, teams, and customer experience around agents. They discuss why Kavak spends as much engineering effort on evals as it does building agents, how its AI sellers outperform its human teams, and an experiment where an AI "CEO" increased profits in one city by 50% in its first month. The conversation also explores what happens to organizational structure when agents do most of the work, why Kavak trains everyone from executives to mechanics to build with AI, and Alejandro's argument that companies looking for incremental AI adoption may be missing the larger opportunity: redesigning the organization itself.   Resources: Follow Alejandro Maza Ayala on X: https://x.com/alehandromz Follow Angela Strange on X: https://x.com/astrange Follow Gabriel Vasquez on X: https://x.com/GEVS94 Stay Updated: Find a16z on YouTube: YouTube Find a16z on X Find a16z on LinkedIn Listen to the a16z Show on Spotify Listen to the a16z Show on Apple Podcasts Follow our host: https://twitter.com/eriktorenberg   Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

  • Your Luxury Car Now Has Ads!

    -4 j

    Your Luxury Car Now Has Ads!

    This week, Marques, Andrew, and David talk about the newest Spiderman x BMW collab that has everyone angry before debating whether the EU should make Apple share some features with Windows computers. Then it's all about CMF and their newest earbuds (modeled by AI models ) and a few quick hits about wearables. It's a long one this week but it was also still a fun one. Enjoy! Links: Reddit - BMW x Spiderman Ford Fathom Verge - Apple to Windows Copy-Paste MacWorld - Fitbit syncs to Apple Health CMF by Nothing Twitter Casio ring watch Keychron Twitter This episode brought to you by: Framer: https://www.framer.com/waveform Shopify: https://www.shopify.com/waveform Follow us on socials: Marques: https://www.threads.net/@mkbhd Andrew: https://www.threads.net/@andrew_manganelli David: https://www.threads.net/@davidimel Adam: https://www.threads.net/@parmesanpapi17 Ellis: https://twitter.com/EllisRovin Waveform Threads: https://www.threads.net/@waveformpodcast Waveform Instagram: https://www.instagram.com/waveformpodcast/?hl=en Waveform TikTok: https://www.tiktok.com/@waveformpodcast Join the Discord: https://discord.gg/mkbhd Intro/Outro music by 20syl: https://bit.ly/2S53xlC Waveform is part of the Vox Media Podcast Network. Learn more about your ad choices. Visit podcastchoices.com/adchoices

  • AI, automation and enshittification

    -23 h

    AI, automation and enshittification

    This special episode of the pod was recorded remotely with tech and science fiction author Cory Doctorow. Cory has recently published a book titled: The Reverse Centaur's Guide to Life After AI, so they start by exploring what a reverse centaur is in this context. Spoiler alert – it’s about the relationship between humans and machines. That takes them down a number of timely rabbit holes regarding what’s in store for us as AI becomes ever more powerful and ubiquitous. They eventually move on to discuss Cory’s previous book, called Enshittification, before concluding with a quick look at the Electronic Frontier Foundation.

  • The playbook for building high talent density teams | Adam Ward, Head of Talent at Cursor

    -1 j

    The playbook for building high talent density teams | Adam Ward, Head of Talent at Cursor

    Adam Ward is the Head of Talent at Cursor, one of the fastest-growing developer tools in history. Before joining Cursor, he founded Growth by Design, an independent recruiting and talent strategy firm that helped build teams at the most ambitious AI and technology companies in the world. Adam has spent more than 20 years building elite, high-talent-density teams across the industry and is widely regarded as one of the most effective and creative recruiters in tech. In our in-depth conversation, we discuss: 1. Inside today’s “tale of two cities” talent market 2. Why the traditional recruiting funnel—what Adam calls the “funnel of doom”—leads to mediocre hires 3. Adam’s three-step playbook: scoping, mapping, and relentless pursuit 4. The worst question to ask when sourcing great talent 5. The rise of the forward deployed engineer—and how to become one 6. The biggest mistake founders make when hiring their first recruiter — Brought to you by: WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more Mercury—Radically different banking, now with Command — Where to find Adam Ward • X: https://x.com/wardadamp • LinkedIn: https://www.linkedin.com/in/adampward — Where to find Lenny: • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ — In this episode, we cover: (00:00) Introduction (03:01) The state of the hiring market (06:41) What roles are trending up and what roles are trending down (11:25) The three-step hiring framework that changes everything (21:58) Getting the right people to actually respond (25:42) Keeping talent a top priority (29:16) Inside Cursor’s recruiting operation (32:19) How to relentlessly pursue top talent (35:42) “Caring is free” (38:33) The conversation most companies underestimate (39:19) Rethinking the atomic unit of a search (40:12) Other recruiting tactics (45:04) What the on-site is really for (49:35) The importance of work trials (53:10) The offer stage (57:53) Negotiations and comp (01:01:50) The details that actually close the deal (01:07:23) What it’s like working with an exec team who values recruiting (01:09:42) Talent density (01:11:18) The first recruiting hire most founders get wrong (01:15:00) Building incredible teams (01:18:12) The time between offer and acceptance (01:19:48) What’s next for the recruiting function (01:21:00) Lightning round — Referenced: • Cursor: https://cursor.com • Trilogy: https://trilogy.com • Brie Wolfson on LinkedIn: https://www.linkedin.com/in/brie-wolfson-17758724 • Inside Cursor: https://colossus.com/article/inside-cursor • SpaceX: https://www.spacex.com • Joe Gebbia’s website: https://joegebbia.com • The rise of Cursor: The $300M ARR AI tool that engineers can’t stop using | Michael Truell (co-founder and CEO): https://www.lennysnewsletter.com/p/the-rise-of-cursor-michael-truell • The Pitt on HBO Max: https://www.hbomax.com/shows/pitt-2024/e6e7bad9-d48d-4434-b334-7c651ffc4bdf • The Bear on Hulu: https://www.hulu.com/series/the-bear-05eb6a8e-90ed-4947-8c0b-e6536cbddd5f • Granola: https://www.granola.ai • Wispr Flow: https://wisprflow.ai • 11 products I love, free for a year—the biggest Product Pass expansion in 2 years: https://www.lennysnewsletter.com/p/productpass-summer2026launch • How to debug a team that isn’t working: the Waterline Model: https://www.lennysnewsletter.com/p/how-to-debug-a-team-that-isnt-working • The high-growth handbook: Molly Graham’s frameworks for leading through chaos, change, and scale: https://www.lennysnewsletter.com/p/the-high-growth-handbook-molly-graham — Recommended books: • Emotional Intelligence: Why It Can Matter More Than IQ: https://www.amazon.com/dp/055338371X • High Growth Handbook: Scaling Startups from 10 to 10,000 People: https://www.amazon.com/High-Growth-Handbook-Elad-Gil/dp/1732265100 • Scaling People: Tactics for Management and Company Building: https://press.stripe.com/scaling-people — Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com. — Lenny may be an investor in the companies discussed. To hear more, visit www.lennysnewsletter.com

  • Disney: The Renaissance and the Empire

    -1 j

    Disney: The Renaissance and the Empire

    In 1984, the Walt Disney Company was worth more dead than alive. Disney Animation — the heart of Walt's famous flywheel — had stagnated for years, bleeding away talent while corporate raiders circled, salivating over offers to sell off the film library to MGM and offload the parks to hotel operators. But what followed instead was the greatest turnaround in media history under Michael Eisner and Frank Wells. Beauty and the Beast. The Lion King. Broadway. Bringing the Disney Vault home on VHS and DVD. And the greatest media acquisition of all time — ESPN. And then... it all almost fell apart. Again. Euro Disney turned into a money pit. Boardroom and executive infighting ran rampant. Animation descended into a dumpster fire. (Remember Chicken Little? Us neither.) Comcast — Comcast!! — tried to steal the company via a hostile takeover. Out of the chaos, a new generation of Disney management emerged under Bob Iger to stage yet another epic comeback with Pixar, Marvel and Lucasfilm, creating the defining media empire of the 21st century…until the tech companies came along. Tune in for the ultimate Acquired thrill ride: Disney, Part II. Sponsors: Many thanks to our fantastic Fall '26 Season partners: SierraSentryWorkOSAnthropicLinks: Sign up for email updates, get our takeaways and research photos from each episode, and vote on future topics!The Official Acquired Meetup on Sept 17th with our friends at Sentry. Join us!The Acquired Disney Part II Companion PDFWorldly Partners' Multi-Decade Disney StudyAll episode sourcesCarve Outs: Warby Parker Transitions Extra ActiveMichael Arndt's Toy Story 3 Story PresentationThe Golden State ValkyriesMore Acquired: Get email updates and vote on future episodes!Join the SlackCheck out the latest swag in the ACQ Merch Store!00:00:00 Start00:00:50 Intro00:05:07 Disney in Chaos (1984)00:11:33 Eisner, Wells, Katzenberg Arrive (1984)00:24:30 Animation Renaissance & CAPS Tech (1989)00:37:33 Flywheel Extensions: Home Video, Retail & Broadway00:54:32 Challenges & ABC/ESPN Acquisition (1994-1995)01:05:55 ESPN: Disney's Accidental Goldmine01:21:26 Eisner's Decline & Save Disney Campaign (2001-2004)01:34:53 Comcast Hostile Takeover Bid (2004)01:41:58 Bob Iger's Vision & Pixar Acquisition (2005-2006)01:52:17 Pixar: From Lucasfilm to Steve Jobs (1979-1995)02:03:11 Toy Story, IPO & Eisner Conflict (1995)02:34:30 Disney Acquires Pixar (2006)02:46:37 Marvel & Lucasfilm Acquisitions (2009-2012)02:58:01 Streaming Pivot: Cord Cutting & BAMTech (2015)03:06:30 The Disney+ Strategy & FOX Acquisition (2017-2019)03:19:01 The Disney+ Launch, COVID, & Chapek's Tenure (2019-2022)03:42:15 Iger's Return, Challenges & Parks Revival (2022-2026)03:50:54 The Business Today: Parks & Streaming Focus03:59:22 Analysis: Disney+ Strategy & The New Media Landscape04:10:01 Analysis: Bull/Bear Cases04:21:20 Quintessence04:24:39 Carve-Outs + Outro ‍Note: Acquired hosts and guests may hold assets discussed in this episode. This podcast is not investment advice, and is intended for informational and entertainment purposes only. You should do your own research and make your own independent decisions when considering any financial transactions.

  • The Clapper: Light work

    19 juil.

    The Clapper: Light work

    Clap on. Clap off. The Clapper may not have been a very good product, but it came with one of the all-time great marketing campaigns. And it might have even been a great, ahead of its time idea. The Verge's David Pierce and Victoria Song are joined by Allison Marsh, a writer and professor at The University of South Carolina, to talk through the Clapper's whole history. It involves lawsuits, broken gadgets, and a pioneering idea about laziness and the smart home. We’re also on video! Check us out on YouTube. Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed. We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11. Learn more about your ad choices. Visit podcastchoices.com/adchoices

  • Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

    8 juil.

    Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

    We’ve been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we’re excited to conclude this series returning to Modal, which has just raised a monster $355M Series C. The cloud was built for developers. But agents are now changing that. The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads. However, agents don’t have that luxury. Now in this new era of agents, everything has to be tighter. They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research: Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time: At the time, Modal was just a teeny little company with a $17M Series A. Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate. In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don’t fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience. We go deep on Modal’s AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal’s capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again. We discuss: * Why Kubernetes wasn’t built for bursty AI workloads * How Modal started as a better runtime before becoming an AI cloud * Why Modal added GPUs before ChatGPT * The shift from developer experience to agent experience * Why observability matters when agents are writing the code * Elastic inference for custom models across audio, video, robotics, and comp bio * GPU snapshotting, cold starts, and why inference workloads are so bursty * Why RL rollouts can require 100,000 sandboxes * DeFlash, speculative decoding, and frontier-level inference performance * Auto Endpoints and making optimized inference easier to deploy * What Modal adds beyond vLLM, SGLang, and raw GPU rental * Modal’s 17-cloud capacity pool and supercloud strategy * Networked sandboxes, sidecars, private IPv6, and RDMA * Serverless multi-node training for post-training and research workloads * Auto-research, model-guided sweeps, and agents launching GPU experiments * Compute strategy, capacity planning, and batch tiers * Why production agents need specialized sandboxes and hard guardrails * Modal’s take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal Bench Akshat Bubna * LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103 * X: https://x.com/akshat_b Modal * Website: https://modal.com Timestamps 00:00:00 Introduction 00:00:39 Modal’s origin and why Kubernetes wasn’t enough 00:04:32 Developer Experience → Agent Experience 00:06:21 Modal’s AI cloud primitives 00:09:14 Sandboxes, agent loops, and proto-Cognition 00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes 00:15:24 DeFlash, speculative decoding, and Auto Endpoints 00:19:59 Production-grade inference beyond raw GPUs 00:22:00 Background agents, Ramp Inspect, and the agent lifecycle 00:24:08 Modal’s 17-cloud supercloud strategy 00:26:40 Networked sandboxes, private IPv6, and RDMA 00:32:48 Multi-node training, post-training, and auto research 00:37:36 Compute strategy, capacity planning, and batch tiers 00:40:55 Open models, real-time AI, and production agent infra 00:43:06 Hard guardrails, managed agents, and specialized sandboxes 00:46:06 Why AI made infrastructure exciting again 00:48:30 Model APIs, differentiated products, and agentic video 00:51:50 CI, coding-agent infra, SDKs, and Modal Bench 00:57:28 Closing Thoughts Transcript Introduction: Modal, Series C, and the Art Party Swyx [00:00:00]: We’re here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C. Akshat [00:00:10]: Thank you. Swyx [00:00:11]: Your party yesterday was amazing. Akshat [00:00:15]: Yeah. Swyx [00:00:15]: From all the photos and all the swag. Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin. Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story. Modal’s Origin: A New Runtime Beyond Kubernetes Akshat [00:00:39]: I first met Eric, who’s the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It’s because you have to run them on Kubernetes. Kubernetes is hard to manage. It’s not built for burstiness and, custom images, Swyx [00:01:03]: Yeah Akshat [00:01:03]: It has a terrible developer experience. Swyx [00:01:05]: And I’ll, I’ll interject Akshat [00:01:06]: Yeah Swyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there’s a bit more of the story there from Spotify and all those things. Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment. Akshat [00:01:26]: Yeah. Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines? Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it’s a very useful primitive in itself. It’s There’s a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would’ve been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out. From Serverless Containers to GPU Workloads Swyx [00:02:19]: Nice. Akshat [00:02:19]: We just didn’t think it would be that big of a deal. Swyx [00:02:22]: Yeah, just like add A100. Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it? Akshat [00:02:28]: Yeah. Primarily it’s just, none of the tooling that was out there was built for, one, a really great developer experience, and also there’s a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there’s just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step. Software-Defined Infrastructure and Decorator-Based DX Swyx [00:03:13]: Yeah. Yeah. Be nice. I don’t know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that. Akshat [00:03:22]: Yeah, the self-provisioning Swyx [00:03:23]: Self-provisioning. Akshat [00:03:24]: Yeah. Swyx [00:03:24]: Yeah. I can’t even remember my own post. Swyx [00:03:26]: And then you put me on the landing page. Akshat [00:03:28]: Yeah. We really like, the term and so we stole it. Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right? Akshat [00:03:37]: Yeah. Swyx [00:03:37]: Was that a big part of the original Akshat [00:03:39]: Yes Swyx [00:03:39]: Story or it was just like a DX layer? Akshat [00:03:41]: That was, really important because we really didn’t want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you’re doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that’s more expressive and dynamic. and so yeah, that was always a very important part. Swyx [00:04:04]: Then the pushback is this is a DSL. Akshat [00:04:07]: Yeah. Swyx [00:04:07]: It’s you’re closed source. I am locked into Modal. Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and su

  • Shame on you

    05/02/2020

    Shame on you

    Mireille and Adam discuss shame as an emotional and experiential construct. We dive into the neural structures involved in processing this emotion as well as the factors and implications of our experience of shame. Shame is a natural response to the threat of vulnerability and perception of oneself as defective or inherently "not enough."

  • The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    3 août

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    Watch the full episode on YouTube: We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3: And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF: Three years ago, inference engineering barely existed as a category. Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem. In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model. The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them. We discuss: * What happens when a 200,000-token request enters an inference system * Cache-aware routing and reusing previously computed KV cache * Why prefill and decode are increasingly handled by different GPUs * When dedicated deployments become cheaper and more reliable than shared APIs * How speculative decoding uses a smaller model to accelerate a larger one * Tool calling, structured outputs, and what LLMs actually do * What it takes to support a new open model on day zero * Grafting Kimi’s vision encoder onto GLM-5.2 * Retrofitting inefficient model layers with components from other architectures * Why models sometimes collapse into repeating the same token * How hardware, kernels, and race conditions create nondeterministic failures * Preserving model fidelity while making inference faster * How quantization errors can cancel each other out * Why inference optimizations still deliver gains of 20%, 100%, and 200% * How optimized serving can make a model up to 10× faster * NVIDIA Dynamo, KV-aware routing, and distributed model serving * Speculative decoding the speculative decoder * Why local AI is about making models less dumb while data-center AI is about making them less slow * Tensor, expert, and pipeline parallelism across GPUs * Hardware-aware model design, auto-tuning, and the case against mega kernels * Rubin and why inference is becoming a systems problem * Whether modern GPUs are evolving into programmable AI ASICs * Why enormous models like Kimi K3 require GB300-class hardware * Why open-source video generation still trails Veo, Kling, and other closed models * The quadratic attention bottleneck behind long-form AI video * Autoregressive video, real-time generation, and compounding quality drift * Why future video systems may combine autoregressive and diffusion architectures * Training for inference and inference for training * Continuous post-training, deployment, evaluation, and improvement loops * How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself * Why faster networking could unlock dramatically faster decoding * Continual learning, KV-cache compaction, and persistent model memory Show Notes * How to build a day-0 API for Kimi K3 * 22580: From GPT2 to Kimi3, Explained Philip Kiely * LinkedIn: https://www.linkedin.com/in/philipkiely * X: https://x.com/philipkiely * Inference Engineering: https://www.baseten.co/inference-engineering/ Ali Taha * LinkedIn: https://www.linkedin.com/in/aliestaha/ * X: https://x.com/waterloointern Timestamps 00:00:00 Introduction and the 200K-Token Prompt 00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling 00:11:26 Launching Production-Ready Open Models 00:19:06 Model Retrofits, Failure Modes, and Nondeterminism 00:28:22 Quantization and Canceling Errors 00:32:15 The Race to 10× Faster Inference 00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI 00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels 01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips 01:10:03 Giant Models and the Limits of GPU Memory 01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation 01:21:47 Audio, Images, and Diffusion Models 01:27:32 Training, Self-Optimizing Models, and Continual Learning 01:40:06 Closing Thoughts Transcript Introduction: Baseten, Waterloo Intern, and Inference Engineering Swyx [00:00:00]: Okay, we’re here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you’ve done, you and I have done before, as well as Ali. Welcome. Ali [00:00:15]: Pleasure to meet you. Swyx [00:00:15]: Waterloo intern. Ali [00:00:16]: Waterloo intern, always. Swyx [00:00:17]: When did you get “Waterloo intern” as a handle? Ali [00:00:19]: As a handle? Oh. Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.” Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer. Philip [00:00:30]: So we have to figure out who’s gonna get the handle. Ali [00:00:33]: Well, I’ll pass the torch over to the next intern. Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad. Ali [00:00:37]: To another Waterloo intern. No, bruh. Philip [00:00:39]: Yeah. Ali [00:00:39]: Intern. Swyx [00:00:40]: Intern, yeah. Ali [00:00:40]: And no. Philip [00:00:41]: You gotta get an intern from Waterloo. Ali [00:00:42]: Yeah, I’ve gotta get an intern from Waterloo. Swyx [00:00:44]: Right. Ali [00:00:44]: But they have to follow the path. Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it’s like whoever Baseten gets from Waterloo. Ali [00:00:48]: Right. Swyx [00:00:49]: Has the title of Waterloo. Ali [00:00:50]: It stays in the ecosystem. Philip [00:00:51]: Exactly. Ali [00:00:52]: Halfway through the internship, you either get it or you’re out. Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle. Ali [00:00:59]: Just say it. Philip [00:00:59]: For everybody. Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you’re an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten’s inference? What’s the process of query through GPU model routing, balancing, all that? What is all the stuff that we don’t think about? Long Context Requests, KV Cache, and Cache-Aware Routing Philip [00:01:26]: With a long query specifically, the first thing that I’m gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it’s gonna be a lot easier for me and a lot cheaper for you. So the first thing that we’re gonna look at is some cache-aware routing, where we’re going to see, we probably have a number of instances, a number of replicas up serving whatever model you’re hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you’re doing two hundred thousand tokens, it’s probably coding or a multi-turn agent or something where you would expect to have that cached. If you don’t, we’re gonna have to send it to a prefill worker. We’ve at least on certain models disaggregated prefill and decode, so you’re going to have one set of GPUs that’s solely going to process the input, create the KV cache, and get you your first token, and then that’s going to be passed over to a separate set of GPUs, which is going to run decode. We’re going to iteratively make those tokens. We’re probably going to have some speculator model in front of that. I’m going to assume that you’re doing coding, and because of that, our speculator model, which assumes you’re doing coding, is gonna have a high draft token acceptance rate. If I’m wrong and you’re asking me to summarize every Harry Potter book, it’s gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?” Swyx [00:03:04]: Except Ba

  • Analyzing OpenAI's Smart Speaker Release

    -3 j

    Analyzing OpenAI's Smart Speaker Release

    In this episode, we take a closer look at the implications of OpenAI's $300 smart speaker on the market. Additionally, we highlight the competitive edge offered by ByteDance's 'Mythos' model.Chapters00:00 OpenAI's Smart Speaker01:57 ByteDance's AI Model05:00 Replit's Self-Driving Code07:58 AI Security Breaches Show Links Get the top 80+ AI Models for $8.99 at AI Box: ⁠⁠https://aibox.ai Get the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Nouveautés

  • Artifacts: Stories from the Emotional History of the Internet
    Documentaire
    Documentaire

    Chaque semaine

  • The Intelligence Shift
    Technologies
    Technologies

    Tous les mois

  • Energy Talks
    Technologies
    Technologies

    Tous les 2 mois

  • Contenu explicite, Andre The Beast Crayton
    Télévision et cinéma
    Télévision et cinéma

    Tous les jours

  • Master Claude Chat, Cowork, Code
    Technologies
    Technologies

    15 mai

  • The Agentic Allocator
    Investissement
    Investissement

    30 juin

  • Artificial Clarity
    Technologies
    Technologies

    16 juin

Actualités technologiques

  • WSJ Tech News Briefing
    Actualités technologiques
    Actualités technologiques

    Tous les jours

  • Bloomberg Tech
    Actualités technologiques
    Actualités technologiques

    Tous les jours

  • Software Engineering Daily
    Actualités technologiques
    Actualités technologiques

    Chaque semaine

  • The AI Engineer Podcast
    Actualités technologiques
    Actualités technologiques

    Tous les jours

  • The WAN Show
    Actualités technologiques
    Actualités technologiques

    Chaque semaine

  • Claude AI
    Actualités technologiques
    Actualités technologiques

    Tous les jours

  • Pharma and BioTech Daily
    Actualités technologiques
    Actualités technologiques

    Tous les jours

  • العربية
  • English (UK)
Choisissez un pays ou une région

Afrique, Moyen‑Orient et Inde

  • Algeria
  • Angola
  • Armenia
  • Azerbaijan
  • Bahrain
  • Benin
  • Botswana
  • Burkina Faso
  • Cameroun
  • Cape Verde
  • Chad
  • Côte d’Ivoire
  • Congo, The Democratic Republic Of The
  • Egypt
  • Eswatini
  • Gabon
  • Gambia
  • Ghana
  • Guinea-Bissau
  • India
  • Iraq
  • Israel
  • Jordan
  • Kenya
  • Kuwait
  • Lebanon
  • Liberia
  • Libya
  • Madagascar
  • Malawi
  • Mali
  • Mauritania
  • Mauritius
  • Morocco
  • Mozambique
  • Namibia
  • Niger (English)
  • Nigeria
  • Oman
  • Qatar
  • Congo, Republic of
  • Rwanda
  • São Tomé and Príncipe
  • Saudi Arabia
  • Senegal
  • Seychelles
  • Sierra Leone
  • South Africa
  • Sri Lanka
  • Tajikistan
  • Tanzania, United Republic Of
  • Tunisia
  • Turkmenistan
  • United Arab Emirates
  • Uganda
  • Yemen
  • Zambia
  • Zimbabwe

Asie‑Pacifique

  • Afghanistan
  • Australia
  • Bhutan
  • Brunei Darussalam
  • Cambodia
  • 中国大陆
  • Fiji
  • 香港
  • Indonesia (English)
  • 日本
  • Kazakhstan
  • 대한민국
  • Kyrgyzstan
  • Lao People's Democratic Republic
  • 澳門
  • Malaysia (English)
  • Maldives
  • Micronesia, Federated States of
  • Mongolia
  • Myanmar
  • Nauru
  • Nepal
  • New Zealand
  • Pakistan
  • Palau
  • Papua New Guinea
  • Philippines
  • Singapore
  • Solomon Islands
  • 台灣
  • Thailand
  • Tonga
  • Turkmenistan
  • Uzbekistan
  • Vanuatu
  • Vietnam

Europe

  • Albania
  • Armenia
  • Österreich
  • Belarus
  • Belgium
  • Bosnia and Herzegovina
  • Bulgaria
  • Croatia
  • Cyprus
  • Czechia
  • Denmark
  • Estonia
  • Finland
  • France (Français)
  • Georgia
  • Deutschland
  • Greece
  • Hungary
  • Iceland
  • Ireland
  • Italia
  • Kosovo
  • Latvia
  • Lithuania
  • Luxembourg (English)
  • Malta
  • Moldova, Republic Of
  • Montenegro
  • Nederland
  • North Macedonia
  • Norway
  • Poland
  • Portugal (Português)
  • Romania
  • Россия
  • Serbia
  • Slovakia
  • Slovenia
  • España
  • Sverige
  • Schweiz
  • Türkiye (English)
  • Ukraine
  • United Kingdom

Amérique latine et Caraïbes

  • Anguilla
  • Antigua and Barbuda
  • Argentina (Español)
  • Bahamas
  • Barbados
  • Belize
  • Bermuda
  • Bolivia (Español)
  • Brasil
  • Virgin Islands, British
  • Cayman Islands
  • Chile (Español)
  • Colombia (Español)
  • Costa Rica (Español)
  • Dominica
  • República Dominicana
  • Ecuador (Español)
  • El Salvador (Español)
  • Grenada
  • Guatemala (Español)
  • Guyana
  • Honduras (Español)
  • Jamaica
  • México
  • Montserrat
  • Nicaragua (Español)
  • Panamá
  • Paraguay (Español)
  • Perú
  • St. Kitts and Nevis
  • Saint Lucia
  • St. Vincent and The Grenadines
  • Suriname
  • Trinidad and Tobago
  • Turks and Caicos
  • Uruguay (English)
  • Venezuela (Español)

États‑Unis et Canada

  • Canada (English)
  • Canada (Français)
  • United States
  • Estados Unidos (Español México)
  • الولايات المتحدة
  • США
  • 美国 (简体中文)
  • États-Unis (Français France)
  • 미국
  • Estados Unidos (Português Brasil)
  • Hoa Kỳ
  • 美國 (繁體中文台灣)

Copyright © 2026 Apple Inc. Tous droits réservés.

  • Conditions générales des services Internet
  • Lecteur Web Apple Podcasts et confidentialité
  • Avertissement concernant les cookies
  • Assistance
  • Remarques

Pour écouter des épisodes au contenu explicite, connectez‑vous.

Apple Podcasts

Recevez les dernières actualités sur cette émission

Connectez‑vous ou inscrivez‑vous pour suivre des émissions, enregistrer des épisodes et recevoir les dernières actualités.

Choisissez un pays ou une région

Afrique, Moyen‑Orient et Inde

  • Algeria
  • Angola
  • Armenia
  • Azerbaijan
  • Bahrain
  • Benin
  • Botswana
  • Burkina Faso
  • Cameroun
  • Cape Verde
  • Chad
  • Côte d’Ivoire
  • Congo, The Democratic Republic Of The
  • Egypt
  • Eswatini
  • Gabon
  • Gambia
  • Ghana
  • Guinea-Bissau
  • India
  • Iraq
  • Israel
  • Jordan
  • Kenya
  • Kuwait
  • Lebanon
  • Liberia
  • Libya
  • Madagascar
  • Malawi
  • Mali
  • Mauritania
  • Mauritius
  • Morocco
  • Mozambique
  • Namibia
  • Niger (English)
  • Nigeria
  • Oman
  • Qatar
  • Congo, Republic of
  • Rwanda
  • São Tomé and Príncipe
  • Saudi Arabia
  • Senegal
  • Seychelles
  • Sierra Leone
  • South Africa
  • Sri Lanka
  • Tajikistan
  • Tanzania, United Republic Of
  • Tunisia
  • Turkmenistan
  • United Arab Emirates
  • Uganda
  • Yemen
  • Zambia
  • Zimbabwe

Asie‑Pacifique

  • Afghanistan
  • Australia
  • Bhutan
  • Brunei Darussalam
  • Cambodia
  • 中国大陆
  • Fiji
  • 香港
  • Indonesia (English)
  • 日本
  • Kazakhstan
  • 대한민국
  • Kyrgyzstan
  • Lao People's Democratic Republic
  • 澳門
  • Malaysia (English)
  • Maldives
  • Micronesia, Federated States of
  • Mongolia
  • Myanmar
  • Nauru
  • Nepal
  • New Zealand
  • Pakistan
  • Palau
  • Papua New Guinea
  • Philippines
  • Singapore
  • Solomon Islands
  • 台灣
  • Thailand
  • Tonga
  • Turkmenistan
  • Uzbekistan
  • Vanuatu
  • Vietnam

Europe

  • Albania
  • Armenia
  • Österreich
  • Belarus
  • Belgium
  • Bosnia and Herzegovina
  • Bulgaria
  • Croatia
  • Cyprus
  • Czechia
  • Denmark
  • Estonia
  • Finland
  • France (Français)
  • Georgia
  • Deutschland
  • Greece
  • Hungary
  • Iceland
  • Ireland
  • Italia
  • Kosovo
  • Latvia
  • Lithuania
  • Luxembourg (English)
  • Malta
  • Moldova, Republic Of
  • Montenegro
  • Nederland
  • North Macedonia
  • Norway
  • Poland
  • Portugal (Português)
  • Romania
  • Россия
  • Serbia
  • Slovakia
  • Slovenia
  • España
  • Sverige
  • Schweiz
  • Türkiye (English)
  • Ukraine
  • United Kingdom

Amérique latine et Caraïbes

  • Anguilla
  • Antigua and Barbuda
  • Argentina (Español)
  • Bahamas
  • Barbados
  • Belize
  • Bermuda
  • Bolivia (Español)
  • Brasil
  • Virgin Islands, British
  • Cayman Islands
  • Chile (Español)
  • Colombia (Español)
  • Costa Rica (Español)
  • Dominica
  • República Dominicana
  • Ecuador (Español)
  • El Salvador (Español)
  • Grenada
  • Guatemala (Español)
  • Guyana
  • Honduras (Español)
  • Jamaica
  • México
  • Montserrat
  • Nicaragua (Español)
  • Panamá
  • Paraguay (Español)
  • Perú
  • St. Kitts and Nevis
  • Saint Lucia
  • St. Vincent and The Grenadines
  • Suriname
  • Trinidad and Tobago
  • Turks and Caicos
  • Uruguay (English)
  • Venezuela (Español)

États‑Unis et Canada

  • Canada (English)
  • Canada (Français)
  • United States
  • Estados Unidos (Español México)
  • الولايات المتحدة
  • США
  • 美国 (简体中文)
  • États-Unis (Français France)
  • 미국
  • Estados Unidos (Português Brasil)
  • Hoa Kỳ
  • 美國 (繁體中文台灣)