AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. −21 h

    FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

    This episode examines FLINT, a proposed hardware/software architecture from Huawei's Zurich research lab (with ETH Zürich and HUST) for closing the gap between multi-terabyte LLM weight sizes and the far smaller on-package memory of today's GPUs. It introduces high bandwidth flash (HBF), an emerging memory tier that stacks 3D NAND dies with through-silicon vias to sit directly beside HBM in the accelerator package, storing read-only model weights while HBM handles fast-changing KV cache and activations. The discussion walks through core NAND flash mechanics — dies, planes, blocks, and pages, along with the punishing asymmetry between microsecond reads and millisecond erases — to explain why naive flash designs stall under refresh operations and static prefetching. It then details how FLINT's burst-buffer controller replaces compiler-driven prefetch hints with real-time demand-based read coalescing, using HBF's built-in page and cache buffers instead of dedicated SRAM. Listeners interested in memory system design, inference hardware economics, and the practical engineering trade-offs of scaling capacity without wasting GPU compute will find this a detailed look at a genuinely emerging technology rather than a shipping product. Sources: 1. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration — Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang, 2026 http://arxiv.org/abs/2608.25062 2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System — Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, Jinho Lee, 2024, HPCA https://scholar.google.com/scholar?q=Smart-Infinity%3A+Fast+Large+Language+Model+Training+using+Near-Storage+Processing+on+a+Real+System 5. InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference — Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang, 2024 https://scholar.google.com/scholar?q=InstInfer%3A+In-Storage+Attention+Offloading+for+Cost-Effective+Long-Context+LLM+Inference 6. DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings — Aayush Gupta, Youngjae Kim, Bhuvan Urgaonkar, 2009, ASPLOS https://scholar.google.com/scholar?q=DFTL%3A+A+Flash+Translation+Layer+Employing+Demand-based+Selective+Caching+of+Page-level+Address+Mappings 7. Design Tradeoffs for SSD Performance — Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D. Davis, Mark Manasse, Rina Panigrahy, 2008, USENIX ATC https://scholar.google.com/scholar?q=Design+Tradeoffs+for+SSD+Performance 8. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs — Matias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh, Damien Le Moal, Gregory R. Ganger, George Amvrosiadis, 2021, USENIX ATC https://scholar.google.com/scholar?q=ZNS%3A+Avoiding+the+Block+Interface+Tax+for+Flash-based+SSDs 9. RAIDR: Retention-Aware Intelligent DRAM Refresh — Jamie Liu, Ben Jaiyen, Richard Veras, Onur Mutlu, 2012, ISCA https://scholar.google.com/scholar?q=RAIDR%3A+Retention-Aware+Intelligent+DRAM+Refresh 10. Threshold Voltage Distribution in MLC NAND Flash Memory: Characterization, Analysis, and Modeling — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, DATE https://scholar.google.com/scholar?q=Threshold+Voltage+Distribution+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Analysis%2C+and+Modeling 11. Error Analysis and Retention-Aware Error Management for NAND Flash Memory — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, Intel Technology Journal https://scholar.google.com/scholar?q=Error+Analysis+and+Retention-Aware+Error+Management+for+NAND+Flash+Memory 12. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — M. Ha, E. Kim, H. Kim, 2026 https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference 13. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — K. Alizadeh, S. I. Mirzadeh, et al., 2024 https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory 14. PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving — A. C. Yüzügüler, J. Zhuang, L. Cavigelli, 2025 https://scholar.google.com/scholar?q=PRESERVE%3A+Prefetching+Model+Weights+and+KV-Cache+in+Distributed+LLM+Serving 15. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs 16. Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory — W. Sun, M. Gao, et al., 2025 https://scholar.google.com/scholar?q=Lincoln%3A+Real-Time+50~100B+LLM+Inference+on+Consumer+Devices+with+LPDDR-Interfaced%2C+Compute-Enabled+Flash+Memory Interactive Visualization: FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

  2. −21 h

    LinearKV: When Exact State Merging Breaks Hybrid Models

    This episode examines LinearKV, a new approach to position-independent caching for hybrid Mamba-attention language models, and a striking result: the mathematically "correct" way to merge cached context — exact algebraic composition of recurrent states — badly degrades one tested model's output quality, while a simpler shortcut using only the most recent cached chunk performs reliably. The discussion contrasts standard prefix caching (used by vLLM's PagedAttention and SGLang's RadixAttention) with position-independent caching, which lets cached chunks be reused regardless of order, and explains why that trick breaks down for linear-recurrent layers like Mamba-2 and Gated DeltaNet, which compress history into a single fixed-size state rather than a token-indexed KV ledger. It traces the core problem to how each cached chunk's recurrent state was built in isolation, making "exact" composition exact only relative to a flawed reference rather than the true full-context computation. Listeners interested in LLM serving infrastructure, caching systems, or the tradeoffs of hybrid architectures will find this a concrete case study in how systems intuitions from full attention can actively mislead when applied to newer recurrent designs. Sources: 1. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs — Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li, 2026 http://arxiv.org/abs/2608.11231 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024 https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs 4. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2025 https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion 5. EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models — Junhao Hu et al., 2025 https://scholar.google.com/scholar?q=EPIC%3A+Efficient+Position-Independent+Context+Caching+for+Serving+Large+Language+Models 6. HYPIC: Accelerating hybrid-attention LLM serving with position-independent caching — Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen, 2026 https://scholar.google.com/scholar?q=HYPIC%3A+Accelerating+hybrid-attention+LLM+serving+with+position-independent+caching 7. Marconi: Prefix caching for the era of hybrid LLMs — Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, Ravi Netravali, 2025 https://scholar.google.com/scholar?q=Marconi%3A+Prefix+caching+for+the+era+of+hybrid+LLMs 8. Gated Delta Networks: Improving Mamba2 with the Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2025 (ICLR) https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+the+Delta+Rule 9. EPIC: Efficient position-independent caching for serving large language models — Junhao Hu, Wenrui Huang, Weidong Wang, et al., 2025 (ICML) https://scholar.google.com/scholar?q=EPIC%3A+Efficient+position-independent+caching+for+serving+large+language+models Interactive Visualization: LinearKV: When Exact State Merging Breaks Hybrid Models

  3. −2 d

    Decoupling KL Direction from Rollout Source in LLM Distillation

    This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question. Sources: 1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026 http://arxiv.org/abs/2605.16826 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015 https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks 4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016 https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks 5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026 https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models 8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025 https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation

  4. −2 d

    On-Policy Distillation: Why a Stronger Teacher Can Backfire

    This episode examines why on-policy distillation (OPD) of large language models can fail catastrophically even when the teacher model is objectively stronger by every benchmark — a 7-billion-parameter teacher completely failed to improve a 1.5-billion-parameter student, while a smaller, weaker teacher succeeded. The discussion traces OPD's mechanics: unlike classic distillation, which trains students on teacher-generated text and suffers from exposure bias, OPD has the student generate its own rollouts and scores them against the teacher's token-level probability distribution via reverse KL divergence, yielding a dense per-token reward without a verifier. The central finding is that teacher-student "overlap ratio" — how much the teacher's likely next tokens actually match the student's own candidate set — determines whether that dense signal teaches anything, meaning benchmark strength and teachability are fundamentally different axes. The paper's three-part structure (phenomenology, mechanism, recipe) is illustrated through a controlled comparison of two similarly-scored Qwen3-4B teacher variants, isolating thinking-pattern compatibility as the real driver of distillation success. Listeners working on post-training or distillation pipelines will find this a direct challenge to the common assumption that upgrading the teacher model is a safe, unconditional improvement. Sources: 1. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding, 2026 http://arxiv.org/abs/2604.13016 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, and colleagues (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28GKD%29 5. Qwen3 Technical Report — An Yang and the Qwen Team (Alibaba), 2025 https://scholar.google.com/scholar?q=Qwen3+Technical+Report 6. On-policy distillation of language models: Learning from self-generated mistakes (MiniLLM) — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28MiniLLM%29 7. Distillation scaling laws — Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb, 2025 https://scholar.google.com/scholar?q=Distillation+scaling+laws 8. On the efficacy of knowledge distillation — Jang Hyun Cho, Bharath Hariharan, 2019 https://scholar.google.com/scholar?q=On+the+efficacy+of+knowledge+distillation 9. Small models struggle to learn from strong reasoners — Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran, 2025 https://scholar.google.com/scholar?q=Small+models+struggle+to+learn+from+strong+reasoners 10. On-policy distillation (Thinking Machines Lab blog) — Kevin Lu and Thinking Machines Lab, 2025 https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab+blog%29 11. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026 https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation Interactive Visualization: On-Policy Distillation: Why a Stronger Teacher Can Backfire

  5. −2 d

    Token Teachability: Rethinking Disagreement in On-Policy Distillation

    This episode examines "Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation," which challenges a core assumption in on-policy knowledge distillation: that raw KL divergence between teacher and student token predictions is a reliable signal for which tokens deserve training focus. The discussion traces the lineage from Hinton's original distillation work through Google DeepMind's on-policy approach, then explains the paper's key insight — large disagreement can mean either a small, actionable correction the student can use, or a "incompatible" mismatch pointing toward options the student assigns near-zero probability, and raw KL can't distinguish the two. Building on this distinction, the authors introduce "token teachability" as a better selection criterion and a method called TA-OPD that trains only on the most teachable tokens, reportedly matching or beating full-dataset training while using just 5% of the tokens. Listeners interested in efficient model training, distillation techniques, or the gap between statistical salience and actual learnability will find the reframing of a decade-old assumption particularly compelling. Sources: 1. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation — Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Hongxia Yang, 2026 http://arxiv.org/abs/2605.26844 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016 https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 5. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 7. Not All Tokens Are What You Need for Pretraining (Rho-1) — Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen, 2024 https://scholar.google.com/scholar?q=Not+All+Tokens+Are+What+You+Need+for+Pretraining+%28Rho-1%29 8. Contrastive Decoding: Open-ended Text Generation as Optimization — Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 9. DistiLLM: Towards Streamlined Distillation for Large Language Models — Jongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young Yun, 2024 https://scholar.google.com/scholar?q=DistiLLM%3A+Towards+Streamlined+Distillation+for+Large+Language+Models 10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 11. TIP: Token Importance in On-Policy Distillation — Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard, 2026 https://scholar.google.com/scholar?q=TIP%3A+Token+Importance+in+On-Policy+Distillation 12. Entropy-Aware On-Policy Distillation of Language Models — Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, Kimin Lee, 2026 https://scholar.google.com/scholar?q=Entropy-Aware+On-Policy+Distillation+of+Language+Models 13. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, et al., 2026 https://scholar.google.com/scholar?q=Rethinking+On-Policy+Distillation+of+Large+Language+Models%3A+Phenomenology%2C+Mechanism%2C+and+Recipe 14. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, et al., 2026 https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning 15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang, Haowei Zhang, et al., 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning Interactive Visualization: Token Teachability: Rethinking Disagreement in On-Policy Distillation

  6. −2 d

    Weak-to-Strong On-Policy Distillation Beats the Teacher

    This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail. Sources: 1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026 http://arxiv.org/abs/2607.26246 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29 5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29 6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision 7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29 8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29 9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024 https://scholar.google.com/scholar?q=Distillation+Scaling+Laws 11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29 Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher

  7. −3 d

    Sleeper Memory Poisoning: When Assistants Remember Lies

    This episode examines "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents," a study showing how attackers can plant fabricated facts into an AI assistant's persistent memory that lie dormant until triggered in an unrelated future conversation. Unlike traditional prompt injection, which manipulates a model's behavior only within a single session, this attack targets the memory-write step itself, allowing a single black-box, universal payload template—refined through an actor-critic search between attacker and critic LLMs—to succeed across arbitrary goals with startlingly high rates (up to 99.8% on GPT-5.5). The discussion breaks down the three-stage pipeline attackers must clear (injection, retrieval, and usage), and highlights a clever technique for maximizing the odds a poisoned memory resurfaces later: rewriting it to boost embedding similarity with plausible future queries while a semantic-consistency judge guards against the rewrite drifting from the original intent. Testing spans 700 document-goal pairs across 15 source types and multiple commercial memory architectures, revealing that whether the model or a separate manager process controls memory writes dramatically changes how exploitable a system is. It's a sobering look at how "memory" — now a default feature across ChatGPT, Claude, Gemini, and agent frameworks like Mem0 — introduces a persistent, hard-to-detect attack surface that outlives the malicious content that created it. Sources: 1. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents — Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz, 2026 http://arxiv.org/abs/2605.15338 2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger et al. (Anthropic), 2024 https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training 3. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, et al., 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 4. Prompt Injection Attack Against LLM-Integrated Applications — Yi Liu et al., 2023 https://scholar.google.com/scholar?q=Prompt+Injection+Attack+Against+LLM-Integrated+Applications 5. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025 https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory 6. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al., 2023 https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models+%28GCG%29 7. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases — Chen, Xiang, Xiao, Song, Li, 2024 (NeurIPS) https://scholar.google.com/scholar?q=AgentPoison%3A+Red-teaming+LLM+Agents+via+Poisoning+Memory+or+Knowledge+Bases 8. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search — Agrawal, Khattab, Potts, 2025 https://scholar.google.com/scholar?q=GEPA%3A+Efficient+Textual+Optimization+via+LLM-based+Reflection+and+Pareto-Efficient+Evolutionary+Search 9. Injection through web agents that fetch pages with hidden HTML instructions — Raghav and Choong, 2026 https://scholar.google.com/scholar?q=Injection+through+web+agents+that+fetch+pages+with+hidden+HTML+instructions 10. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers — Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger, 2026 https://scholar.google.com/scholar?q=The+Trigger+in+the+Haystack%3A+Extracting+and+Reconstructing+LLM+Backdoor+Triggers Interactive Visualization: Sleeper Memory Poisoning: When Assistants Remember Lies

  8. −3 d

    Why AI Systems Don't Learn After Deployment

    This episode examines a paper by Emmanuel Dupoux, Yann LeCun, and Jitendra Malik arguing that deployed AI models learn nothing after training, unlike a toddler who continuously experiments through action, observation, imitation, and inquiry. The discussion breaks down the paper's core distinction between System A (passive, observation-based statistical learning like self-supervised training) and System B (action-based reinforcement learning through feedback), and explains why neither alone can produce autonomous intelligence. It then covers the paper's proposed fix, System M, an orchestrator modeled on software-defined networking that monitors low-bandwidth "meta-state" signals like prediction error and confidence to dynamically route between learning systems, automating what human MLOps engineers currently do by hand. The conversation also connects this framework to LeCun's 2022 autonomous machine intelligence proposal and the ongoing debate sparked by Silver and Sutton's "Era of Experience" critique about AI hitting a data wall. Listeners interested in the architecture of autonomous learning and what's actually missing between today's static models and genuinely adaptive intelligence will find the systems-level framing illuminating. Sources: 1. Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science — Emmanuel Dupoux, Yann LeCun, Jitendra Malik, 2026 http://arxiv.org/abs/2603.15381 2. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022 https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence 3. Welcome to the Era of Experience — David Silver, Richard Sutton, 2025 https://scholar.google.com/scholar?q=Welcome+to+the+Era+of+Experience 4. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) — Julian Schrittwieser et al., 2020 https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model+%28MuZero%29 5. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran et al. (incl. LeCun), 2025 https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning 6. Coordination Among Neural Modules Through a Shared Global Workspace — Anirudh Goyal, Aniket Didolkar, et al., 2022 https://scholar.google.com/scholar?q=Coordination+Among+Neural+Modules+Through+a+Shared+Global+Workspace 7. Embodied AI Agents: Modeling the World — Pascale Fung, Emmanuel Dupoux, Jitendra Malik, et al., 2025 https://scholar.google.com/scholar?q=Embodied+AI+Agents%3A+Modeling+the+World Interactive Visualization: Why AI Systems Don't Learn After Deployment

Om

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

Du kanske också gillar