Distributed Shampoo: Making Second-Order Optimization Practical at Scale

This episode explores a distributed-systems paper from Meta, Mila/University of Montreal, and NVIDIA researchers that tackles a long-standing tradeoff in neural network training: second-order-style optimizers like Shampoo converge better than the ubiquitous diagonal methods (AdaGrad, RMSProp, Adam), but were assumed too computationally expensive to use at scale. The discussion traces the lineage from diagonal AdaGrad through the impractical "full-matrix" AdaGrad to Shampoo's key innovation — approximating each layer's preconditioner as a Kronecker product of two much smaller matrices, drawing a parallel to Martens and Grosse's independently-derived KFAC method. The hosts debate whether Shampoo counts as a true second-order method or something distinct rooted in online convex optimization theory, and highlight the paper's headline systems result: a distributed PyTorch implementation that keeps the wall-clock overhead of this matrix-based optimizer to roughly 10% per step. Listeners interested in the practical engineering behind making theoretically superior optimizers actually usable at billion-parameter scale will find the breakdown of the memory and compute tradeoffs especially compelling. Sources: 1. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, Michael Rabbat, 2023 http://arxiv.org/abs/2309.06497 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer (Google), 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 7. On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models — Rohan Anil, Sandra Gadanho, Da Huang, et al., 2022 https://scholar.google.com/scholar?q=On+the+Factory+Floor%3A+ML+Engineering+for+Industrial-Scale+Ads+Recommendation+Models 8. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost — Noam Shazeer, Mitchell Stern, 2018 https://scholar.google.com/scholar?q=Adafactor%3A+Adaptive+Learning+Rates+with+Sublinear+Memory+Cost 9. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam Interactive Visualization: Distributed Shampoo: Making Second-Order Optimization Practical at Scale