Turing Post

Turing Post

Hi, I’m Ksenia, founder of Turing Post. On this channel, I talk to the people shaping AI and pay attention to the ideas, shifts, and details others might miss. Inference is my interview show with innovators, builders, founders, and thinkers moving AI forward. Attention Span is where I slow down on what deserves a closer look: the signals, questions, and stories hiding between the headlines. Subscribe for the unusual takes. And always stay curious!

  1. 21h ago

    Does AI Understand the Machine It Runs On? | Inside OpenAI

    What if a model finds an optimization that surprises the engineers who have spent years working on that system? What, exactly, has it understood? I brought that question to OpenAI’s Phil Tillet and Matt Ferrari, whose work involves making AI cheaper and more accessible. They’re increasingly doing that work with the models themselves. Matt talks about research ideas his team used to dismiss because the engineering would be too complicated. Now they can give a model years of earlier research and ask it to explore what might work. They’re using models to help improve the smaller models behind speculative decoding, including the training process itself. So I ask whether the model is also helping choose the ideas, and how much that expands what they’re willing to try. I also ask something familiar to anyone who uses these systems: why does the same model sometimes feel different? Phil explains that even changing the order of floating-point calculations can introduce differences in its behavior. That puts a very concrete problem behind our conversation about understanding: an optimization can make the system faster and still change something you wanted to preserve. We get into how they catch those changes, and what happens when the failure is something nobody thought to test for. *We talk about:* When models became useful for engineering decisions. Why OpenAI’s models needed more control than Triton gave them. What happens inside the system after you send a request. How model-assisted kernel improvements helped cut Sol’s serving costs by 20%. Letting models investigate bugs independently, and deciding when to step in. Why successfully optimizing something can still be a waste of effort. Whether models need an internal representation of how a computer system behaves. How AI assistance opens up experiments that engineers previously couldn’t justify attempting. *Chapters:*   *Follow on*: https://www.turingpost.com/ *Did you like the episode? You know the drill:*  📌 Subscribe here and here (https://www.turingpost.com/subscribe) for more conversations with the builders shaping real-world AI.  💬 Leave a comment 👍 Like it  🫶 Thank you for watching and sharing! *Guests:*  Philippe (Phil) Tillet created Triton, a programming language that makes efficient GPU programming more accessible. He joined OpenAI as an intern in 2019, before it had an API or a product, and spent years improving training efficiency. His interests extend from compilers and kernels to Bertrand Russell and philosophy of mind. Matthew (Matt) Ferrari works on inference efficiency at OpenAI, across request routing, load balancing, debugging and speculative decoding. His fascination with optimization began in school, when GPU programming changed his understanding of how fast an algorithm could run. Today, he brings that curiosity to the entire system serving a model. #openai #inference #optimization

  2. 6d ago

    NVIDIA’s $12.9B Plan to Rule Open-Source AI

    NVIDIA was once so hostile to open source that Linus Torvalds gave the company the finger. Today, it maintains open Linux modules, releases hundreds of models and datasets, and has reportedly agreed to buy Hugging Face for $12.9 billion. WHAT?! The change makes sense once we examine what NVIDIA learned from nearly dying with NV1, spending years searching for CUDA’s market, and watching researchers discover deep learning on gaming GPUs. In this episode, we follow that strategy from NV1 and CUDA to NVIDIA’s reported $12.9 billion acquisition of Hugging Face, and ask whether the company has found the most profitable model for open-source AI. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to explain the technical and business choices shaping AI. Links: NVIDIA FY2026 Form 10-K https://www.sec.gov/Archives/edgar/data/1045810/000104581026000021/nvda-20260125.htm  Interview with Spencer Huang (Nvidia) https://www.youtube.com/watch?v=NEv9EnD7JVU&t=1231 Interview with Clem Delangue (Hugging Face) https://www.youtube.com/watch?v=DfJV722V1WY  NVIDIA Open Source https://opensource.nvidia.com/en-us NVIDIA on Hugging Face https://huggingface.co/nvidia Reuters on the reported $12.9B agreement https://www.reuters.com/technology/nvidia-talks-acquire-hugging-face-13-billion-deal-business-insider-reports-2026-08-27/ NVIDIA Open GPU Kernel Modules https://github.com/NVIDIA/open-gpu-kernel-modules Turing Post’s history of computer vision and AlexNet https://www.turingpost.com/p/cvhistory6 #NVIDIA #OpenSourceAI #HuggingFace #AI #GitHub

  3. Aug 31

    Fei-Fei Li, LeCun, Hassabis: What Do They Mean by “World Model”?

    Demis Hassabis, Yann LeCun, Fei-Fei Li – they all talk about “building a world model.” Some of them are dedicating their professional lives to it! But do they mean the same? So before joining the World Models workshop at Chicago Booth, I wanted to answer a basic question: what do researchers mean by a world model, and how many different ideas are sitting under this name? World models are absolutely fascinating area of research with its GPT moment still in the nearest future.  This episode is based on the current research and provides a comprehensive overview of three broad approaches: generating future observations, predicting inside learned representations such as JEPA, and learning only what a planner needs to make decisions. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the best way to understand a model is to change the background to purple and see what breaks. *Links:*  Demis Hassabis on world models https://www.youtube.com/watch?v=sZaM6MadDZU Yann LeCun on world models https://www.youtube.com/watch?v=8sS9UJzb_t4 Fei-Fei Li on large world models https://www.youtube.com/watch?v=pNYVckbCFuk Beyond LLMs: JEPA and the Road to AGI – the main milestones so far https://www.youtube.com/watch?v=z0fh0SY3VWc stable-worldmodel https://github.com/galilai-group/stable-worldmodel/issues/153 VideoPhy-2, a benchmark https://arxiv.org/pdf/2503.06800  Physion-Eval https://arxiv.org/html/2603.19607v1  What Is JEPA? LeCun Architecture & World Models https://www.turingpost.com/p/jepa  #WorldModels #AI #MachineLearning #YannLeCun #FeiFeiLi #DemisHassabis #JEPA #PhysicalAI #TuringPost #AttentionSpan

  4. Aug 28

    OpenCode vs. OpenRouter: The Fight Over Your AI Models

    OpenCode began as an open-source coding agent. Now it is selling model access, negotiating directly with suppliers and preparing to reserve its own GPU capacity. That puts it on a collision course with OpenRouter, the model marketplace Stripe has agreed to acquire for a reported $8 billion. This episode follows this new shift in the industry and what Ox Alpha showed about the value of distribution: the company controlling the workflow may influence which models win long before a developer opens the model menu. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 Interviews: @realturingpost Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com #OpenCode #OpenRouter #AIAgents #CodingAgents #AIInfrastructure Sources and further reading OpenRouter is joining Stripe https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/  OpenCode https://opencode.ai/ Ox Alpha, Explained Without the Hype https://www.youtube.com/watch?v=tN8xiPoareo&t=16s OpenCode Zen https://opencode.ai/docs/zen/ GLM-5.3-Flash, formerly Ox Alpha, usage data https://opencode.ai/data/zhipuai/glm-5.3-flash Dax Raad on OpenCode’s direction https://x.com/thdxr/status/2093161006226612377 Dax Raad on inference economics https://x.com/thdxr/status/2093161006226612377 Dax Raad on OpenCode’s buying power https://x.com/thdxr/status/2092844520119345160 Jay V on OpenCode’s token volume https://x.com/snowmaker/status/2080667637861011924

  5. Aug 25

    Ox Alpha, Explained Without the Hype

    An anonymous model called Ox Alpha appeared on OpenRouter and OpenCode on August 20 with a million-token context window, video input, and a price of zero. Within four days it had processed tens of trillions of tokens, and the internet had spent those same four days trying to work out who built it. In this episode:  how you fingerprint a model you know nothing about,  why the evidence points at Z.ai's unreleased multimodal GLM,  what the 113-task benchmark runs really show versus the viral 80 percent,  the three contradictory data policies governing your prompts,  and the thought I keep coming back to – that the platform a model launches on is becoming as decisive as the lab that trained it. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 Interviews: @realturingpost Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com Sources and further reading  Ox Alpha vs GLM-5.3 on OpenRouter: https://openrouter.ai/compare/stealth/ox-alpha/z-ai/glm-5.3  Ox Alpha on OpenCode https://opencode.ai/data/unknown/ox-alpha OpenCode Zen documentation https://dev.opencode.ai/docs/zen OpenRouter Stealth Model Terms https://openrouter.ai/terms/stealth The Tokenizer Is a Fingerprint by Joseph Elstner https://isimplifyme.com/whitepapers/the-tokenizer-is-a-fingerprint DeepSWE result https://x.com/winkey_h/status/2090814178810306874/photo/1  58.4% run, MatchaOnMuffins/oxalpha https://github.com/MatchaOnMuffins/oxalpha/blob/main/README.md 64.6% run, jyeric/ox-alpha-deepswe https://github.com/jyeric/ox-alpha-deepswe/blob/main/README.md Community fingerprinting summary: https://cellcog.ai/blog/what-is-ox-alpha/  Prediction market on the reveal: https://manifold.markets/Sketchy/who-is-behind-ox-alpha-the-mysterio  #OxAlpha #OpenRouter #OpenCode #GLM #AIcoding #stealthmodel

  6. Aug 25

    Etched Explained: The $21B AI Chip Startup Challenging NVIDIA

    Etched raised $1 billion in 26 days. Its valuation jumped from $10.3 billion to $21 billion. The second round was led by Jane Street after it tested Etched’s hardware and installed the first rack in its own data center. So what did Jane Street see? Etched began with Sohu, a Transformer-only ASIC that promised more than 500,000 tokens per second on Llama 70B. By 2026, Sohu and that claim had disappeared. Etched now sells a complete inference cluster and says it can run Transformers, MoEs, and even Mamba. We explain how that shift is possible, what Low Voltage Inference and Cluster Scale Memory actually mean, and how this still tiny company can hurt giant NVIDIA. And the question I want you to keep from this episode: The GPU once found the winning middle ground between flexibility and specialization. Has Etched found the next one? *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the decisive question is how much flexibility we are still willing to pay for. *Sources:* Etched, From Zero to One Etched, Accelerating Inference and Frontier Inference Clusters Etched’s 2026 architecture-agnostic and Mamba claims Reuters on the $21 billion financing and Jane Street deployment The Wall Street Journal on Etched’s team and NVIDIA recruiting TechCrunch on the original Transformer-only Sohu pitch Etched patent on model-specific ASIC compilation and configurable execution Mamba-2 and Structured State Space Duality Jane Street on its machine-learning infrastructure Jane Street on microsecond-scale performance engineering CoreWeave and Jane Street’s $6 billion cloud agreement The founders on Etched’s supply-chain choices NVIDIA on Vera Rubin and Groq 3 LPX NVIDIA Q1 FY2027 results Taalas on model-specific silicon #AttentionSpan #Etched #JaneStreet #AIChips #AIInference #NVIDIA #Semiconductors #Mamba #TuringPost

  7. Aug 25

    Why DeepSeek Harness Is The End Of Coding Agents as We Know Them

    DeepSeek just open-sourced Harness – it can write its own missing tools while it runs, then cleanly remove them. 149k GitHub stars in four days. An 88-page paper underneath. It’s open, easy to install and it claims that *everything is a plugin.*  What does it mean? We unpack that plus we discuss why DeepSeek Harness is not another Claude Code clone, but the moment the fixed coding agent starts to die. We also look at the history of computing (Smalltalk, Unix, Codd) to ask whether “everything is a plugin” can do what “everything is an object” and “everything is a file” once did. The question I want you to think about: once an agent can recompose itself, what exactly is the product anymore? *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the decisive move is engineering the layer everyone else treated as packaging. *Links* - DeepSeek Harness repository and installation: https://github.com/deepseek-ai/deepseek-harness - Cordis repository: https://github.com/cordiverse/cordis - Cordis paper: https://github.com/cordiverse/paper/blob/main/paper.pdf - Koishi introduction, Touhou name origin, community, and plugin history: https://koishi.chat/en-US/manual/introduction - Koishi repository: https://github.com/koishijs/koishi - SmallTalk History https://computerhistory.org/blog/introducing-the-smalltalk-zoo-48-years-of-smalltalk-history-at-chm/ - Sholto Douglas post: https://x.com/_sholtodouglas/status/2088463770318516734 #AttentionSpan #DeepSeekHarness #DeepSeek #AIAgents #CodingAgents #EverythingIsAPlugin #OpenSource #Cordis #AgentArchitecture #TuringPost

About

Hi, I’m Ksenia, founder of Turing Post. On this channel, I talk to the people shaping AI and pay attention to the ideas, shifts, and details others might miss. Inference is my interview show with innovators, builders, founders, and thinkers moving AI forward. Attention Span is where I slow down on what deserves a closer look: the signals, questions, and stories hiding between the headlines. Subscribe for the unusual takes. And always stay curious!

You Might Also Like