AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 2d ago

    Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

    This episode examines "Is Grep All You Need? How Agent Harnesses Reshape Agentic Search," which tests whether simple regex-based retrieval can outperform vector search inside agentic pipelines like Chronos, Claude Code, and Codex CLI. The hosts dig into how retrieval mode interacts with harness architecture, delivery method (inline vs. programmatic), and backbone model choice, finding that inline grep beats inline vector search across every harness-model pairing tested — with gaps as wide as twenty points and swings as large as switching harnesses entirely. A striking case shows the same model scoring 93.1% on one harness but only 76.7% on another, suggesting orchestration and prompt construction matter as much as the retrieval algorithm itself. The discussion also surfaces a counterintuitive twist: forcing an agent to read retrieved results from a file instead of getting them dumped inline can nearly halve accuracy, even with identical underlying search. Listeners interested in RAG, agent design, or LLM evaluation methodology will find the paper's tangled-but-honest approach to measuring real deployed systems a useful corrective to cleaner but less realistic ablation studies. Sources: 1. Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sahil Sen, Akhil Kasturi, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 http://arxiv.org/abs/2605.15184 2. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022 https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models 3. WebGPT: Browser-assisted question-answering with human feedback — Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al. (OpenAI), 2021 https://scholar.google.com/scholar?q=WebGPT%3A+Browser-assisted+question-answering+with+human+feedback 4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks 5. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Di Wu, Hongwei Wang, Wenhao Yu, et al., 2024 https://scholar.google.com/scholar?q=LongMemEval%3A+Benchmarking+Chat+Assistants+on+Long-Term+Interactive+Memory 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond 7. Dense Passage Retrieval for Open-Domain Question Answering — Vladimir Karpukhin, Barlas Oğuz, Sewon Min, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Dense+Passage+Retrieval+for+Open-Domain+Question+Answering 8. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, et al., 2023 https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts 9. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Thibault Formal, Benjamin Piwowarski, Stéphane Clinchant, 2021 https://scholar.google.com/scholar?q=SPLADE%3A+Sparse+Lexical+and+Expansion+Model+for+First+Stage+Ranking 10. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models — Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych, 2021 https://scholar.google.com/scholar?q=BEIR%3A+A+Heterogenous+Benchmark+for+Zero-shot+Evaluation+of+Information+Retrieval+Models 11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 12. Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory — Sahil Sen, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 https://scholar.google.com/scholar?q=Chronos%3A+Temporal-Aware+Conversational+Agents+with+Structured+Event+Retrieval+for+Long-Term+Memory Interactive Visualization: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

  2. 3d ago

    AI-Generated Text and the Death of the Open Web

    This episode examines a study analyzing the growth of AI-generated content across the open web, drawing on 33 monthly samples from the Internet Archive's Wayback Machine between August 2022 and May 2025. It highlights the paper's central finding that AI-generated or AI-assisted content on newly published websites rose from zero before ChatGPT's launch to roughly 35 percent by mid-2025, and explores how the authors transform "Dead Internet Theory" from internet folklore into six testable hypotheses — including semantic contraction, truth decay, positivity shift, epistemic islands, entropy dilution, and stylistic monoculture. The discussion covers the methodology behind sampling a representative slice of the internet, including logarithmic downsampling and stratification across time, MIME type, and domain to avoid bias toward heavily-crawled sites. It also connects the findings to the concept of model collapse, framing the 35 percent figure as empirical evidence for a previously theoretical concern about AI models training on their own synthetic output. Listeners interested in web ecosystem health, LLM training data quality, or the intersection of internet culture and rigorous data science will find the episode's blend of meme-to-metric translation particularly compelling. Sources: 1. The Impact of AI-Generated Text on the Internet — Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek, 2026 http://arxiv.org/abs/2604.26965 2. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature — Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn, 2023 https://scholar.google.com/scholar?q=DetectGPT%3A+Zero-Shot+Machine-Generated+Text+Detection+using+Probability+Curvature 3. Can AI-Generated Text be Reliably Detected? — Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, Soheil Feizi, 2023 https://scholar.google.com/scholar?q=Can+AI-Generated+Text+be+Reliably+Detected%3F 4. A Watermark for Large Language Models — John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023 https://scholar.google.com/scholar?q=A+Watermark+for+Large+Language+Models 5. GLTR: Statistical Detection and Visualization of Generated Text — Sebastian Gehrmann, Hendrik Strobelt, Alexander M. Rush, 2019 https://scholar.google.com/scholar?q=GLTR%3A+Statistical+Detection+and+Visualization+of+Generated+Text 6. The Curse of Recursion: Training on Generated Data Makes Models Forget — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, Ross Anderson, 2024 (Nature, updated from 2023 preprint) https://scholar.google.com/scholar?q=The+Curse+of+Recursion%3A+Training+on+Generated+Data+Makes+Models+Forget 7. Is the Internet Dead? Evaluating Claims about Web Homogenization from AI-Generated Content (representative of the broader 2024-2025 measurement literature, e.g. Muzumdar et al. and related environmental-scanning studies of Dead Internet Theory discourse) — Various (this cluster of work is cited in the paper as Muzumdar et al., 2025 and similar), 2025 https://scholar.google.com/scholar?q=Is+the+Internet+Dead%3F+Evaluating+Claims+about+Web+Homogenization+from+AI-Generated+Content+%28representative+of+the+broader+2024-2025+measurement+literature%2C+e.g.+Muzumdar+et+al.+and+related+environmental-scanning+studies+of+Dead+Internet+Theory+discourse%29 8. Studies on bot/inauthentic-content prevalence on specific platforms (e.g., La Cava et al. 2025 on social media, Matatov et al. 2024 on platform-specific AI content) — Lucio La Cava et al.; Jonathan Matatov et al. (representative platform-specific studies cited in the paper's related-work section), 2024-2025 https://scholar.google.com/scholar?q=Studies+on+bot%2Finauthentic-content+prevalence+on+specific+platforms+%28e.g.%2C+La+Cava+et+al.+2025+on+social+media%2C+Matatov+et+al.+2024+on+platform-specific+AI+content%29 9. Public Trust and Perceptions of Artificial Intelligence (Ipsos / Reuters Institute Digital News Report and Edelman Trust Barometer AI-focused editions) — Ipsos (various); Reuters Institute for the Study of Journalism (Nic Newman et al.); Edelman Trust Barometer team, 2023-2025 (recurring annual) https://scholar.google.com/scholar?q=Public+Trust+and+Perceptions+of+Artificial+Intelligence+%28Ipsos+%2F+Reuters+Institute+Digital+News+Report+and+Edelman+Trust+Barometer+AI-focused+editions%29 10. Americans' Views of Artificial Intelligence (Pew Research Center recurring survey series) — Pew Research Center (Alec Tyson, Emma Kikuchi, and colleagues), 2023-2025 (recurring) https://scholar.google.com/scholar?q=Americans%27+Views+of+Artificial+Intelligence+%28Pew+Research+Center+recurring+survey+series%29 11. The Perception Gap: Comparing Public Beliefs about Misinformation to Empirical Prevalence Estimates (representative of the risk-perception vs. measured-prevalence literature this paper's framing descends from, e.g. work following Duffy et al. and general misperception-of-misinformation-prevalence studies) — Andrew Guess, colleagues in the misinformation-prevalence research cluster (representative of this line, distinct from the AI-specific surveys above), 2019-2023 (foundational misinformation-perception literature) https://scholar.google.com/scholar?q=The+Perception+Gap%3A+Comparing+Public+Beliefs+about+Misinformation+to+Empirical+Prevalence+Estimates+%28representative+of+the+risk-perception+vs.+measured-prevalence+literature+this+paper%27s+framing+descends+from%2C+e.g.+work+following+Duffy+et+al.+and+general+misperception-of-misinformation-prevalence+studies%29 12. Documenting the English Colossal Clean Crawled Corpus (C4) — Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner, 2021 https://scholar.google.com/scholar?q=Documenting+the+English+Colossal+Clean+Crawled+Corpus+%28C4%29 13. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only — Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay, 2023 https://scholar.google.com/scholar?q=The+RefinedWeb+Dataset+for+Falcon+LLM%3A+Outperforming+Curated+Corpora+with+Web+Data%2C+and+Web+Data+Only 14. Quantifying Memorization Across Neural Language Models / broader Common Crawl representativeness and bias studies (e.g., work on Common Crawl's domain and language skew) — Various (Common Crawl bias/representativeness literature, e.g. work by Luccioni & Viviano on Common Crawl content quality, and follow-on studies), 2021-2023 https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models+%2F+broader+Common+Crawl+representativeness+and+bias+studies+%28e.g.%2C+work+on+Common+Crawl%27s+domain+and+language+skew%29 15. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, et al. (Allen Institute for AI), 2024 https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research 16. AI models collapse when trained on recursively generated data — I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal, 2024 https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data 17. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks — V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, 2025 https://scholar.google.com/scholar?q=Can+AI-generated+text+be+reliably+detected%3F+Stress+testing+AI+text+detectors+under+various+attacks 18. RAID: A shared benchmark for robust evaluation of machine-generated text detectors — L. Dugan, A. Hwang, F. Trhlik, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, C. Callison-Burch, 2024 https://scholar.google.com/scholar?q=RAID%3A+A+shared+benchmark+for+robust+evaluation+of+machine-generated+text+detectors 19. When incentives backfire, data stops being human — S. Santy, P. Bhattacharya, M. H. Ribeiro, K. Allen, S. Oh, 2025 https://scholar.google.com/scholar?q=When+incentives+backfire%2C+data+stops+being+human 20. Longitudinal sampling of URLs from the Wayback Machine — K. Garg, S. Alam, D. Ayala, M. Graham, M. C. Weigle, M. L. Nelson, 2025 https://scholar.google.com/scholar?q=Longitudinal+sampling+of+URLs+from+the+Wayback+Machine 21. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity — J. Zhang, S. Yu, D. Chong, A. Sicilia, M. R. Tomz, C. D. Manning, W. Shi, 2025 https://scholar.google.com/scholar?q=Verbalized+sampling%3A+How+to+mitigate+mode+collapse+and+unlock+LLM+diversity Interactive Visualization: AI-Generated Text and the Death of the Open Web

  3. 3d ago

    Approaching Shannon Bound: Lossless LLM Weight Compression

    This episode explores "Approaching Shannon Bound with Lossless LLM Weight Compression," which argues that model weights stored in formats like bf16 carry far less real information than their bit-width implies—entropy measurements across six models and seven numeric formats show gaps of several bits per weight that can be recovered without any change to the underlying values. The discussion covers why memory capacity and bandwidth, not raw compute, are the real bottleneck in GPU inference, and why generic compressors like gzip fail on IEEE-754 floating point data. The hosts dig into Asymmetric Numeral Systems (ANS) as the key engineering breakthrough, since it decodes fast enough and in a tile-parallel enough fashion to run inside a live GPU kernel without becoming a new bottleneck itself. Listeners interested in the intersection of information theory and practical LLM serving will find the walkthrough of how lossless compression differs fundamentally from quantization methods like int4 or AWQ particularly compelling. Sources: 1. Approaching Shannon Bound with Lossless LLM Weight Compression — Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He, 2026 http://arxiv.org/abs/2606.15789 2. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float — Tianyi Zhang, Yang Sui, Shaochen (Henry) Zhong, et al., 2025 https://scholar.google.com/scholar?q=70%25+Size%2C+100%25+Accuracy%3A+Lossless+LLM+Compression+for+Efficient+GPU+Inference+via+Dynamic-Length+Float 3. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks — Yongchang Hao, Yanshuai Cao, Lili Mou, 2024 https://scholar.google.com/scholar?q=NeuZip%3A+Memory-Efficient+Training+and+Inference+with+Dynamic+Compression+of+Neural+Networks 4. ZipNN: Lossless Compression for AI Models — Moshik Hershcovitch, Andrew Wood, Leshem Choshen, et al. (IBM Research), 2024 https://scholar.google.com/scholar?q=ZipNN%3A+Lossless+Compression+for+AI+Models 5. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016 https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding 6. Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding — Jarek Duda, 2013 https://scholar.google.com/scholar?q=Asymmetric+Numeral+Systems%3A+Entropy+Coding+Combining+Speed+of+Huffman+Coding+with+Compression+Rate+of+Arithmetic+Coding 7. The Use of Asymmetric Numeral Systems as an Accurate Replacement for Huffman Coding — Jarek Duda, Khalid Tahboub, Neeraj J. Gadgil, Edward J. Delp, 2015 https://scholar.google.com/scholar?q=The+Use+of+Asymmetric+Numeral+Systems+as+an+Accurate+Replacement+for+Huffman+Coding 8. Variational Image Compression with a Scale Hyperprior — Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, Nick Johnston, 2018 https://scholar.google.com/scholar?q=Variational+Image+Compression+with+a+Scale+Hyperprior 9. Zstandard Compression and the application/zstd Media Type (RFC 8878) — Yann Collet, Murray Kucherawy (eds.), 2020 https://scholar.google.com/scholar?q=Zstandard+Compression+and+the+application%2Fzstd+Media+Type+%28RFC+8878%29 10. GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMs — Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, Tuo Zhao, 2024 https://scholar.google.com/scholar?q=GEAR%3A+An+Efficient+KV+Cache+Compression+Recipe+for+Near-Lossless+Generative+Inference+of+LLMs 11. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters 12. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024 https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving 13. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are All You Need — M. Davies, N. Crago, K. Sankaralingam, C. Kozyrakis, 2025 https://scholar.google.com/scholar?q=Efficient+LLM+Inference%3A+Bandwidth%2C+Compute%2C+Synchronization%2C+and+Capacity+are+All+You+Need Interactive Visualization: Approaching Shannon Bound: Lossless LLM Weight Compression

  4. 3d ago

    Decomposing Speedups Across Runtime, Kernel, and Quantization

    This episode dissects a common but misleading claim in LLM serving benchmarks: that swapping to a faster stack alone explains a headline speedup number. It walks through the mechanics separating prefill's compute-bound math from decode's memory-bandwidth-bound token generation, then explains how continuous batching and PagedAttention keep GPUs saturated, and how GPTQ quantization paired with the Marlin kernel avoids costly dequantization overhead. By introducing a matched FP16 intermediate stack, the paper cleanly splits a reported speedup into a runtime factor and a kernel-plus-quantization factor that multiply back to the observed total. The headline finding is striking: a controlled, apples-to-apples comparison yields a 2.58x speedup dominated by runtime (68%), while a more common operational comparison — pitting a batch-capped baseline against a fully loaded modern stack — inflates that to 10.62x, with runtime's share climbing to 86%. Listeners get a rare, rigorous look at how much of the "free lunch" in inference speedups is real architecture versus benchmark framing. Sources: 1. Decomposing Speedups Across Runtime, Kernel, and Quantization https://arxiv.org/pdf/2607.11368 2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 3. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale 4. SqueezeLLM: Dense-and-Sparse Quantization — Sehoon Kim, Coleman Hooper, Amir Gholami, et al., 2023 https://scholar.google.com/scholar?q=SqueezeLLM%3A+Dense-and-Sparse+Quantization 5. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, et al., 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML Interactive Visualization: Decomposing Speedups Across Runtime, Kernel, and Quantization

  5. 4d ago

    Distributed Shampoo: Making Second-Order Optimization Practical at Scale

    This episode explores a distributed-systems paper from Meta, Mila/University of Montreal, and NVIDIA researchers that tackles a long-standing tradeoff in neural network training: second-order-style optimizers like Shampoo converge better than the ubiquitous diagonal methods (AdaGrad, RMSProp, Adam), but were assumed too computationally expensive to use at scale. The discussion traces the lineage from diagonal AdaGrad through the impractical "full-matrix" AdaGrad to Shampoo's key innovation — approximating each layer's preconditioner as a Kronecker product of two much smaller matrices, drawing a parallel to Martens and Grosse's independently-derived KFAC method. The hosts debate whether Shampoo counts as a true second-order method or something distinct rooted in online convex optimization theory, and highlight the paper's headline systems result: a distributed PyTorch implementation that keeps the wall-clock overhead of this matrix-based optimizer to roughly 10% per step. Listeners interested in the practical engineering behind making theoretically superior optimizers actually usable at billion-parameter scale will find the breakdown of the memory and compute tradeoffs especially compelling. Sources: 1. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, Michael Rabbat, 2023 http://arxiv.org/abs/2309.06497 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer (Google), 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 7. On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models — Rohan Anil, Sandra Gadanho, Da Huang, et al., 2022 https://scholar.google.com/scholar?q=On+the+Factory+Floor%3A+ML+Engineering+for+Industrial-Scale+Ads+Recommendation+Models 8. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost — Noam Shazeer, Mitchell Stern, 2018 https://scholar.google.com/scholar?q=Adafactor%3A+Adaptive+Learning+Rates+with+Sublinear+Memory+Cost 9. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam Interactive Visualization: Distributed Shampoo: Making Second-Order Optimization Practical at Scale

  6. 4d ago

    SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

    This episode explores SOAP, a new optimizer from a Harvard/Kempner Institute team that fuses Shampoo's second-order preconditioning with Adam's update mechanics. The hosts trace the lineage from Adagrad's mathematically ideal but computationally infeasible full preconditioner matrix, through Adam's cheap diagonal approximation, to Shampoo's middle-ground Kronecker-product approach using two smaller per-dimension preconditioners. The core theoretical result discussed is a proof that Shampoo run with the one-half power is mathematically equivalent to running Adafactor inside the eigenbasis Shampoo's own preconditioner defines — which motivates simply swapping in full Adam within that same rotated basis, adding just one new hyperparameter (preconditioning frequency) over standard AdamW. The discussion highlights the paper's striking efficiency claims — over 40% fewer training iterations and 35% less wall-clock time versus AdamW, and roughly 20% better than Shampoo itself — while noting these numbers deserve scrutiny given real-world context like Shampoo's AlgoPerf benchmark win and its use in training Gemini 1.5 Flash. Listeners interested in the mechanics behind large-scale training efficiency will get a clear breakdown of why optimizer choice translates directly into cluster-scale compute costs and calendar time. Sources: 1. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Mujin Kwun, Itai Shapira, David Brandfonbrener, Lucas Janson, Sham Kakade, 2024 http://arxiv.org/abs/2409.11321 2. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 3. 4-bit Shampoo for Memory-Efficient Network Training — Sike Wang, Jia Li, Pan Zhou, Hua Huang, 2024 https://scholar.google.com/scholar?q=4-bit+Shampoo+for+Memory-Efficient+Network+Training 4. Combining axes preconditioners through Kronecker approximation for deep learning — Sai Surya Duvvuri, Fnu Devvrit, Rohan Anil, Cho-Jui Hsieh, Inderjit S. Dhillon, 2024 https://scholar.google.com/scholar?q=Combining+axes+preconditioners+through+Kronecker+approximation+for+deep+learning 5. No train no gain: Revisiting efficient training algorithms for transformer-based language models — Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner, 2023 https://scholar.google.com/scholar?q=No+train+no+gain%3A+Revisiting+efficient+training+algorithms+for+transformer-based+language+models Interactive Visualization: SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

  7. 4d ago

    Test-Time Adaptation Through Entropy Minimization

    This episode explores Tent, a 2021 method for adapting a frozen classifier to shifted test data using only entropy minimization on unlabeled target inputs — no retraining, no labels, and no access to the original source dataset. It contrasts this "fully test-time adaptation" setting against classical domain adaptation, which still requires the source data on hand during adjustment, a constraint that's often impractical for vendors shipping models under privacy or bandwidth limits. The discussion digs into the mechanism: Tent re-estimates BatchNorm statistics on incoming test batches and tunes only the tiny per-channel scale-and-shift parameters (under 1% of the network), repurposing existing training infrastructure for adaptation. The hosts also interrogate the core intuition — that confident predictions tend to be correct — pressing on whether that assumption holds when a model's decision boundaries are already unreliable, without fully resolving the tension before turning to results. Listeners interested in low-cost deployment fixes for distribution shift, or skeptical of self-referential confidence-based methods, will find the back-and-forth pushback especially engaging. Sources: 1. Tent: Fully Test-time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2020 http://arxiv.org/abs/2006.10726 2. Semi-Supervised Learning by Entropy Minimization — Yves Grandvalet, Yoshua Bengio, 2004 https://scholar.google.com/scholar?q=Semi-Supervised+Learning+by+Entropy+Minimization 3. A DIRT-T Approach to Unsupervised Domain Adaptation — Rui Shu, Hung Bui, Hirokazu Narui, Stefano Ermon, 2018 https://scholar.google.com/scholar?q=A+DIRT-T+Approach+to+Unsupervised+Domain+Adaptation 4. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020 https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts 5. Unsupervised Domain Adaptation by Backpropagation — Yaroslav Ganin, Victor Lempitsky, 2015 https://scholar.google.com/scholar?q=Unsupervised+Domain+Adaptation+by+Backpropagation 6. Learning from Synthetic Data: Addressing Domain Shift for Semantic Segmentation — Swami Sankaranarayanan, Yogesh Balaji, Arpit Jain, Ser Nam Lim, Rama Chellappa, 2018 https://scholar.google.com/scholar?q=Learning+from+Synthetic+Data%3A+Addressing+Domain+Shift+for+Semantic+Segmentation 7. Revisiting Batch Normalization For Practical Domain Adaptation — Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, Xiaodi Hou, 2016 https://scholar.google.com/scholar?q=Revisiting+Batch+Normalization+For+Practical+Domain+Adaptation 8. CyCADA: Cycle-Consistent Adversarial Domain Adaptation — Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei A. Efros, Trevor Darrell, 2018 https://scholar.google.com/scholar?q=CyCADA%3A+Cycle-Consistent+Adversarial+Domain+Adaptation 9. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift — Sergey Ioffe, Christian Szegedy, 2015 https://scholar.google.com/scholar?q=Batch+Normalization%3A+Accelerating+Deep+Network+Training+by+Reducing+Internal+Covariate+Shift 10. Evaluating Prediction-Time Batch Normalization for Robustness to Covariate Shift — Zachary Nado, Shreyas Padhy, D. Sculley, Alexander D'Amour, Balaji Lakshminarayanan, Jasper Snoek, 2020 https://scholar.google.com/scholar?q=Evaluating+Prediction-Time+Batch+Normalization+for+Robustness+to+Covariate+Shift 11. Improving Robustness Against Common Corruptions by Covariate Shift Adaptation — Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, Matthias Bethge, 2020 https://scholar.google.com/scholar?q=Improving+Robustness+Against+Common+Corruptions+by+Covariate+Shift+Adaptation 12. Test-Time Training for Out-of-Distribution Generalization — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2019 https://scholar.google.com/scholar?q=Test-Time+Training+for+Out-of-Distribution+Generalization 13. Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation (SHOT) — Jian Liang, Dapeng Hu, Jiashi Feng, 2020 https://scholar.google.com/scholar?q=Do+We+Really+Need+to+Access+the+Source+Data%3F+Source+Hypothesis+Transfer+for+Unsupervised+Domain+Adaptation+%28SHOT%29 14. Do CIFAR-10 Classifiers Generalize to CIFAR-10? / Do ImageNet Classifiers Generalize to ImageNet? — Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, Vaishaal Shankar, 2018 / 2019 https://scholar.google.com/scholar?q=Do+CIFAR-10+Classifiers+Generalize+to+CIFAR-10%3F+%2F+Do+ImageNet+Classifiers+Generalize+to+ImageNet%3F 15. A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions (ANT) — Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann, Matthias Bethge, Wieland Brendel, 2020 https://scholar.google.com/scholar?q=A+Simple+Way+to+Make+Neural+Networks+Robust+Against+Diverse+Image+Corruptions+%28ANT%29 Interactive Visualization: Test-Time Adaptation Through Entropy Minimization

  8. 5d ago

    Preconditioned Optimization Without the Full Matrix Cost

    This episode explores Shampoo, a 2018 optimization algorithm from Google Brain and Princeton that brings second-order curvature information to neural network training without the prohibitive cost of a full preconditioner. The discussion traces the lineage from Newton's method through AdaGrad's diagonal and full-matrix variants, explaining how Shampoo exploits the natural tensor structure of neural network weights — keeping separate small preconditioners per dimension and combining them via Kronecker products rather than materializing an impossibly large matrix. The hosts debate how much weight to put on the algorithm's convergence proofs, which rely on online convex optimization theory even though real network training is highly non-convex, concluding that the theory offers a sanity check rather than a guarantee and that empirical performance does the real work of justification. Along the way, they place Shampoo alongside K-FAC as one of the major structure-aware curvature approximations in the optimization literature. Listeners interested in the tradeoffs between cheap first-order methods and expensive second-order ones will find a clear walkthrough of how Shampoo threads that needle. Sources: 1. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 http://arxiv.org/abs/1802.09568 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Scalable Second Order Optimization for Deep Learning (Distributed Shampoo) — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2021 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning+%28Distributed+Shampoo%29 4. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 5. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 6. Adam: A Method for Stochastic Optimization — Diederik Kingma, Jimmy Ba, 2015 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 7. Decoupled Weight Decay Regularization — Ilya Loshchilov, Frank Hutter, 2017 https://scholar.google.com/scholar?q=Decoupled+Weight+Decay+Regularization 8. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 9. A Stochastic Quasi-Newton Method for Large-Scale Optimization — Richard Byrd, Samantha Hansen, Jorge Nocedal, Yoran Singer, 2016 https://scholar.google.com/scholar?q=A+Stochastic+Quasi-Newton+Method+for+Large-Scale+Optimization 10. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 11. Online Convex Programming and Generalized Infinitesimal Gradient Ascent — Martin Zinkevich, 2003 https://scholar.google.com/scholar?q=Online+Convex+Programming+and+Generalized+Infinitesimal+Gradient+Ascent 12. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012 https://scholar.google.com/scholar?q=Online+Learning+and+Online+Convex+Optimization 13. Attention Is All You Need — A. Vaswani, N. Shazeer, N. Parmar, et al., 2017 https://scholar.google.com/scholar?q=Attention+Is+All+You+Need 14. A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization — V. Gupta, T. Koren, Y. Singer, 2017 https://scholar.google.com/scholar?q=A+Unified+Approach+to+Adaptive+Regularization+in+Online+and+Stochastic+Optimization 15. Understanding Deep Learning Requires Rethinking Generalization — C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, 2017 https://scholar.google.com/scholar?q=Understanding+Deep+Learning+Requires+Rethinking+Generalization Interactive Visualization: Preconditioned Optimization Without the Full Matrix Cost

Ratings & Reviews

3.7
out of 5
3 Ratings

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like