AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 1 day ago

    Decoupling KL Direction from Rollout Source in LLM Distillation

    This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question. Sources: 1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026 http://arxiv.org/abs/2605.16826 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015 https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks 4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016 https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks 5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026 https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models 8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025 https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation

  2. 1 day ago

    On-Policy Distillation: Why a Stronger Teacher Can Backfire

    This episode examines why on-policy distillation (OPD) of large language models can fail catastrophically even when the teacher model is objectively stronger by every benchmark — a 7-billion-parameter teacher completely failed to improve a 1.5-billion-parameter student, while a smaller, weaker teacher succeeded. The discussion traces OPD's mechanics: unlike classic distillation, which trains students on teacher-generated text and suffers from exposure bias, OPD has the student generate its own rollouts and scores them against the teacher's token-level probability distribution via reverse KL divergence, yielding a dense per-token reward without a verifier. The central finding is that teacher-student "overlap ratio" — how much the teacher's likely next tokens actually match the student's own candidate set — determines whether that dense signal teaches anything, meaning benchmark strength and teachability are fundamentally different axes. The paper's three-part structure (phenomenology, mechanism, recipe) is illustrated through a controlled comparison of two similarly-scored Qwen3-4B teacher variants, isolating thinking-pattern compatibility as the real driver of distillation success. Listeners working on post-training or distillation pipelines will find this a direct challenge to the common assumption that upgrading the teacher model is a safe, unconditional improvement. Sources: 1. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding, 2026 http://arxiv.org/abs/2604.13016 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, and colleagues (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28GKD%29 5. Qwen3 Technical Report — An Yang and the Qwen Team (Alibaba), 2025 https://scholar.google.com/scholar?q=Qwen3+Technical+Report 6. On-policy distillation of language models: Learning from self-generated mistakes (MiniLLM) — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28MiniLLM%29 7. Distillation scaling laws — Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb, 2025 https://scholar.google.com/scholar?q=Distillation+scaling+laws 8. On the efficacy of knowledge distillation — Jang Hyun Cho, Bharath Hariharan, 2019 https://scholar.google.com/scholar?q=On+the+efficacy+of+knowledge+distillation 9. Small models struggle to learn from strong reasoners — Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran, 2025 https://scholar.google.com/scholar?q=Small+models+struggle+to+learn+from+strong+reasoners 10. On-policy distillation (Thinking Machines Lab blog) — Kevin Lu and Thinking Machines Lab, 2025 https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab+blog%29 11. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026 https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation Interactive Visualization: On-Policy Distillation: Why a Stronger Teacher Can Backfire

  3. 1 day ago

    Token Teachability: Rethinking Disagreement in On-Policy Distillation

    This episode examines "Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation," which challenges a core assumption in on-policy knowledge distillation: that raw KL divergence between teacher and student token predictions is a reliable signal for which tokens deserve training focus. The discussion traces the lineage from Hinton's original distillation work through Google DeepMind's on-policy approach, then explains the paper's key insight — large disagreement can mean either a small, actionable correction the student can use, or a "incompatible" mismatch pointing toward options the student assigns near-zero probability, and raw KL can't distinguish the two. Building on this distinction, the authors introduce "token teachability" as a better selection criterion and a method called TA-OPD that trains only on the most teachable tokens, reportedly matching or beating full-dataset training while using just 5% of the tokens. Listeners interested in efficient model training, distillation techniques, or the gap between statistical salience and actual learnability will find the reframing of a decade-old assumption particularly compelling. Sources: 1. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation — Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Hongxia Yang, 2026 http://arxiv.org/abs/2605.26844 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016 https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 5. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 7. Not All Tokens Are What You Need for Pretraining (Rho-1) — Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen, 2024 https://scholar.google.com/scholar?q=Not+All+Tokens+Are+What+You+Need+for+Pretraining+%28Rho-1%29 8. Contrastive Decoding: Open-ended Text Generation as Optimization — Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 9. DistiLLM: Towards Streamlined Distillation for Large Language Models — Jongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young Yun, 2024 https://scholar.google.com/scholar?q=DistiLLM%3A+Towards+Streamlined+Distillation+for+Large+Language+Models 10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 11. TIP: Token Importance in On-Policy Distillation — Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard, 2026 https://scholar.google.com/scholar?q=TIP%3A+Token+Importance+in+On-Policy+Distillation 12. Entropy-Aware On-Policy Distillation of Language Models — Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, Kimin Lee, 2026 https://scholar.google.com/scholar?q=Entropy-Aware+On-Policy+Distillation+of+Language+Models 13. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, et al., 2026 https://scholar.google.com/scholar?q=Rethinking+On-Policy+Distillation+of+Large+Language+Models%3A+Phenomenology%2C+Mechanism%2C+and+Recipe 14. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, et al., 2026 https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning 15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang, Haowei Zhang, et al., 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning Interactive Visualization: Token Teachability: Rethinking Disagreement in On-Policy Distillation

  4. 1 day ago

    Weak-to-Strong On-Policy Distillation Beats the Teacher

    This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail. Sources: 1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026 http://arxiv.org/abs/2607.26246 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29 5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29 6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision 7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29 8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29 9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024 https://scholar.google.com/scholar?q=Distillation+Scaling+Laws 11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29 Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher

  5. 2 days ago

    Sleeper Memory Poisoning: When Assistants Remember Lies

    This episode examines "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents," a study showing how attackers can plant fabricated facts into an AI assistant's persistent memory that lie dormant until triggered in an unrelated future conversation. Unlike traditional prompt injection, which manipulates a model's behavior only within a single session, this attack targets the memory-write step itself, allowing a single black-box, universal payload template—refined through an actor-critic search between attacker and critic LLMs—to succeed across arbitrary goals with startlingly high rates (up to 99.8% on GPT-5.5). The discussion breaks down the three-stage pipeline attackers must clear (injection, retrieval, and usage), and highlights a clever technique for maximizing the odds a poisoned memory resurfaces later: rewriting it to boost embedding similarity with plausible future queries while a semantic-consistency judge guards against the rewrite drifting from the original intent. Testing spans 700 document-goal pairs across 15 source types and multiple commercial memory architectures, revealing that whether the model or a separate manager process controls memory writes dramatically changes how exploitable a system is. It's a sobering look at how "memory" — now a default feature across ChatGPT, Claude, Gemini, and agent frameworks like Mem0 — introduces a persistent, hard-to-detect attack surface that outlives the malicious content that created it. Sources: 1. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents — Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz, 2026 http://arxiv.org/abs/2605.15338 2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger et al. (Anthropic), 2024 https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training 3. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, et al., 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 4. Prompt Injection Attack Against LLM-Integrated Applications — Yi Liu et al., 2023 https://scholar.google.com/scholar?q=Prompt+Injection+Attack+Against+LLM-Integrated+Applications 5. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025 https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory 6. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al., 2023 https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models+%28GCG%29 7. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases — Chen, Xiang, Xiao, Song, Li, 2024 (NeurIPS) https://scholar.google.com/scholar?q=AgentPoison%3A+Red-teaming+LLM+Agents+via+Poisoning+Memory+or+Knowledge+Bases 8. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search — Agrawal, Khattab, Potts, 2025 https://scholar.google.com/scholar?q=GEPA%3A+Efficient+Textual+Optimization+via+LLM-based+Reflection+and+Pareto-Efficient+Evolutionary+Search 9. Injection through web agents that fetch pages with hidden HTML instructions — Raghav and Choong, 2026 https://scholar.google.com/scholar?q=Injection+through+web+agents+that+fetch+pages+with+hidden+HTML+instructions 10. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers — Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger, 2026 https://scholar.google.com/scholar?q=The+Trigger+in+the+Haystack%3A+Extracting+and+Reconstructing+LLM+Backdoor+Triggers Interactive Visualization: Sleeper Memory Poisoning: When Assistants Remember Lies

  6. 2 days ago

    Why AI Systems Don't Learn After Deployment

    This episode examines a paper by Emmanuel Dupoux, Yann LeCun, and Jitendra Malik arguing that deployed AI models learn nothing after training, unlike a toddler who continuously experiments through action, observation, imitation, and inquiry. The discussion breaks down the paper's core distinction between System A (passive, observation-based statistical learning like self-supervised training) and System B (action-based reinforcement learning through feedback), and explains why neither alone can produce autonomous intelligence. It then covers the paper's proposed fix, System M, an orchestrator modeled on software-defined networking that monitors low-bandwidth "meta-state" signals like prediction error and confidence to dynamically route between learning systems, automating what human MLOps engineers currently do by hand. The conversation also connects this framework to LeCun's 2022 autonomous machine intelligence proposal and the ongoing debate sparked by Silver and Sutton's "Era of Experience" critique about AI hitting a data wall. Listeners interested in the architecture of autonomous learning and what's actually missing between today's static models and genuinely adaptive intelligence will find the systems-level framing illuminating. Sources: 1. Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science — Emmanuel Dupoux, Yann LeCun, Jitendra Malik, 2026 http://arxiv.org/abs/2603.15381 2. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022 https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence 3. Welcome to the Era of Experience — David Silver, Richard Sutton, 2025 https://scholar.google.com/scholar?q=Welcome+to+the+Era+of+Experience 4. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) — Julian Schrittwieser et al., 2020 https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model+%28MuZero%29 5. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran et al. (incl. LeCun), 2025 https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning 6. Coordination Among Neural Modules Through a Shared Global Workspace — Anirudh Goyal, Aniket Didolkar, et al., 2022 https://scholar.google.com/scholar?q=Coordination+Among+Neural+Modules+Through+a+Shared+Global+Workspace 7. Embodied AI Agents: Modeling the World — Pascale Fung, Emmanuel Dupoux, Jitendra Malik, et al., 2025 https://scholar.google.com/scholar?q=Embodied+AI+Agents%3A+Modeling+the+World Interactive Visualization: Why AI Systems Don't Learn After Deployment

  7. 3 days ago

    Second-Order Optimization Meets Runtime Scheduling at Scale

    This episode examines "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," a May 2026 Oxford paper introducing Asteria, a runtime system rather than a new optimizer. The discussion covers why curvature-aware methods like Shampoo and SOAP have never displaced AdamW despite converging in fewer steps, tracing the lineage from K-FAC through Distributed Shampoo to SOAP and explaining the Kronecker-factorization tricks that make tracking curvature tractable at all. The hosts unpack the paper's "three physical walls" framework — a vertical capacity wall from single-GPU memory limits, an overlap disruption wall where cubic-cost matrix operations stall compute-communication overlap, and a global consensus wall from synchronous full-state updates across mismatched network speeds — and debate whether reengineering the plumbing around an unchanged optimizer counts as a genuine research contribution. Listeners interested in distributed training infrastructure, optimizer design trade-offs, or the gap between algorithmic elegance and practical deployability will find the back-and-forth over real benchmark numbers (96 seconds versus 1.5 seconds per step) especially grounded. Sources: 1. Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training — Yishun Lu, Junhao Zhang, Zeyu Yang, Wes Armour, 2026 http://arxiv.org/abs/2605.16184 2. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 6. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — H.-J. M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat, 2023 https://scholar.google.com/scholar?q=A+Distributed+Data-Parallel+PyTorch+Implementation+of+the+Distributed+Shampoo+Optimizer+for+Training+Neural+Networks+At-Scale 7. Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading — A. Maurya, J. Ye, M. M. Rafique, F. Cappello, B. Nicolae, 2024 https://scholar.google.com/scholar?q=Deep+Optimizer+States%3A+Towards+Scalable+Training+of+Transformer+Models+Using+Interleaved+Offloading 8. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, Y. He, 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 9. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization — W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, R. B. Grosse, 2026 https://scholar.google.com/scholar?q=Understanding+and+Improving+Shampoo+and+SOAP+via+Kullback-Leibler+Minimization 10. Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training — Y. Lu, W. Armour, 2026 https://scholar.google.com/scholar?q=Beyond+the+Mean%3A+Fisher-Orthogonal+Projection+for+Natural+Gradient+Descent+in+Large+Batch+Training Interactive Visualization: Second-Order Optimization Meets Runtime Scheduling at Scale

  8. 4 days ago

    Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup

    This episode digs into a paper testing whether hand-written PTX assembly beats NVIDIA's WMMA API for Tensor Core GEMM kernels on an L4 GPU, finding that the answer flips depending on numeric precision rather than holding as a universal rule. The discussion covers the hardware distinction between Tensor Cores and regular CUDA cores, and contrasts the convenience of WMMA against the finer control PTX offers through instructions like cp.async, ldmatrix, and mma.sync. A key thread traces why this matters in practice: quantized LLM serving at INT8 or INT4 shifts kernels from compute-bound to memory-bound, making the precision-dependent payoff of hand-tuned PTX directly relevant to running open-weight models cheaply. The episode also addresses the methodological choice to test on a single GPU, arguing that isolating precision and instruction-set effects requires holding hardware constant rather than spreading across devices. Listeners get a concrete framework for deciding when the extra engineering effort of writing raw PTX is worth it versus when it's wasted work. Sources: 1. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 — Matt J. Borowski, Blazej Osinski, 2026 http://arxiv.org/abs/2608.10103 2. NVIDIA Tensor Core Programmability, Performance & Precision — Stefano Markidis, Steven W. D. Chien, Erwin Laure, Ivy B. Peng, Jeffrey S. Vetter, 2018 https://scholar.google.com/scholar?q=NVIDIA+Tensor+Core+Programmability%2C+Performance+%26+Precision 3. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2018 https://scholar.google.com/scholar?q=Dissecting+the+NVIDIA+Volta+GPU+Architecture+via+Microbenchmarking 4. CUTLASS: CUDA Templates for Linear Algebra Subroutines — Andrew Kerr, Duane Merrill, Julien Demouth, John Tran (NVIDIA), with ongoing project contributors, 2018 (initial release, actively maintained since) https://scholar.google.com/scholar?q=CUTLASS%3A+CUDA+Templates+for+Linear+Algebra+Subroutines 5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 6. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019 https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations 7. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020 (OSDI) https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning 8. Dissecting the Ampere GPU Architecture via Microbenchmarking — Wei Sun, Ang Li, Tong Geng, Sander Stuijk, Henk Corporaal, 2022 https://scholar.google.com/scholar?q=Dissecting+the+Ampere+GPU+Architecture+via+Microbenchmarking 9. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 (OSDI) https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning 10. Understanding Latency Hiding on GPUs — V. Volkov, 2016 https://scholar.google.com/scholar?q=Understanding+Latency+Hiding+on+GPUs 11. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — J. Lin et al., 2024 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 12. FlashInfer: Kernel Library for LLM Serving — Z. Ye et al., 2024 https://scholar.google.com/scholar?q=FlashInfer%3A+Kernel+Library+for+LLM+Serving

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like