This episode examines "Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation," which challenges a core assumption in on-policy knowledge distillation: that raw KL divergence between teacher and student token predictions is a reliable signal for which tokens deserve training focus. The discussion traces the lineage from Hinton's original distillation work through Google DeepMind's on-policy approach, then explains the paper's key insight — large disagreement can mean either a small, actionable correction the student can use, or a "incompatible" mismatch pointing toward options the student assigns near-zero probability, and raw KL can't distinguish the two. Building on this distinction, the authors introduce "token teachability" as a better selection criterion and a method called TA-OPD that trains only on the most teachable tokens, reportedly matching or beating full-dataset training while using just 5% of the tokens. Listeners interested in efficient model training, distillation techniques, or the gap between statistical salience and actual learnability will find the reframing of a decade-old assumption particularly compelling. Sources: 1. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation — Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Hongxia Yang, 2026 http://arxiv.org/abs/2605.26844 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016 https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 5. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 7. Not All Tokens Are What You Need for Pretraining (Rho-1) — Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen, 2024 https://scholar.google.com/scholar?q=Not+All+Tokens+Are+What+You+Need+for+Pretraining+%28Rho-1%29 8. Contrastive Decoding: Open-ended Text Generation as Optimization — Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 9. DistiLLM: Towards Streamlined Distillation for Large Language Models — Jongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young Yun, 2024 https://scholar.google.com/scholar?q=DistiLLM%3A+Towards+Streamlined+Distillation+for+Large+Language+Models 10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 11. TIP: Token Importance in On-Policy Distillation — Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard, 2026 https://scholar.google.com/scholar?q=TIP%3A+Token+Importance+in+On-Policy+Distillation 12. Entropy-Aware On-Policy Distillation of Language Models — Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, Kimin Lee, 2026 https://scholar.google.com/scholar?q=Entropy-Aware+On-Policy+Distillation+of+Language+Models 13. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, et al., 2026 https://scholar.google.com/scholar?q=Rethinking+On-Policy+Distillation+of+Large+Language+Models%3A+Phenomenology%2C+Mechanism%2C+and+Recipe 14. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, et al., 2026 https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning 15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang, Haowei Zhang, et al., 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning Interactive Visualization: Token Teachability: Rethinking Disagreement in On-Policy Distillation