This episode examines "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," a May 2026 Oxford paper introducing Asteria, a runtime system rather than a new optimizer. The discussion covers why curvature-aware methods like Shampoo and SOAP have never displaced AdamW despite converging in fewer steps, tracing the lineage from K-FAC through Distributed Shampoo to SOAP and explaining the Kronecker-factorization tricks that make tracking curvature tractable at all. The hosts unpack the paper's "three physical walls" framework — a vertical capacity wall from single-GPU memory limits, an overlap disruption wall where cubic-cost matrix operations stall compute-communication overlap, and a global consensus wall from synchronous full-state updates across mismatched network speeds — and debate whether reengineering the plumbing around an unchanged optimizer counts as a genuine research contribution. Listeners interested in distributed training infrastructure, optimizer design trade-offs, or the gap between algorithmic elegance and practical deployability will find the back-and-forth over real benchmark numbers (96 seconds versus 1.5 seconds per step) especially grounded. Sources: 1. Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training — Yishun Lu, Junhao Zhang, Zeyu Yang, Wes Armour, 2026 http://arxiv.org/abs/2605.16184 2. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 6. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — H.-J. M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat, 2023 https://scholar.google.com/scholar?q=A+Distributed+Data-Parallel+PyTorch+Implementation+of+the+Distributed+Shampoo+Optimizer+for+Training+Neural+Networks+At-Scale 7. Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading — A. Maurya, J. Ye, M. M. Rafique, F. Cappello, B. Nicolae, 2024 https://scholar.google.com/scholar?q=Deep+Optimizer+States%3A+Towards+Scalable+Training+of+Transformer+Models+Using+Interleaved+Offloading 8. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, Y. He, 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 9. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization — W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, R. B. Grosse, 2026 https://scholar.google.com/scholar?q=Understanding+and+Improving+Shampoo+and+SOAP+via+Kullback-Leibler+Minimization 10. Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training — Y. Lu, W. Armour, 2026 https://scholar.google.com/scholar?q=Beyond+the+Mean%3A+Fisher-Orthogonal+Projection+for+Natural+Gradient+Descent+in+Large+Batch+Training Interactive Visualization: Second-Order Optimization Meets Runtime Scheduling at Scale
Information
- Show
- FrequencyUpdated Daily
- Published26 August 2026 at 12:00 am UTC
- RatingClean
