AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 1d ago

    MegaTrain: Training 120B-Parameter Models on a Single GPU

    This episode examines MegaTrain, a method for full-precision training of 100-billion-parameter-plus language models on a single H200 GPU paired with 1.5 terabytes of host RAM, from a Notre Dame and Lehigh University team. The discussion centers on why memory, not compute, is the real bottleneck for most researchers, citing a survey showing only two of 167 surveyed U.S. universities average more than one H100 per student, while post-training work like instruction tuning and alignment increasingly demands full parameter and optimizer states without full pretraining-scale hardware. The hosts walk through the GPU memory hierarchy — from on-chip SRAM through HBM, host DDR5, and NVMe — and the 12-bytes-per-parameter cost of Adam optimizer state that makes a 70B model require 840 gigabytes of persistent storage. They contrast MegaTrain's approach with prior offloading systems like ZeRO-Offload and ZeRO-Infinity, highlighting the key architectural inversion: host memory becomes the authoritative store for all parameters and optimizer state, while GPU HBM is reduced to a transient scratchpad streaming one layer at a time across PCIe. Listeners interested in democratizing large-model training on constrained hardware will find the systems-level tradeoffs and pointed debate over whether this is genuinely novel or a repackaging of known offloading techniques particularly engaging. Sources: 1. MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU — Zhengqing Yuan, Hanchi Sun, Lichao Sun, Yanfang Ye, 2026 http://arxiv.org/abs/2604.05091 2. vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design — Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, Stephen W. Keckler, 2016 https://scholar.google.com/scholar?q=vDNN%3A+Virtualized+Deep+Neural+Networks+for+Scalable%2C+Memory-Efficient+Neural+Network+Design 3. ZeRO-Offload: Democratizing Billion-Scale Model Training — Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, Yuxiong He, 2021 https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training 4. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He, 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 5. JAX: composable transformations of Python+NumPy programs — James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, Skye Wanderman-Milne, 2018 https://scholar.google.com/scholar?q=JAX%3A+composable+transformations+of+Python%2BNumPy+programs 6. Chainer: A Next-Generation Open Source Framework for Deep Learning — Seiya Tokui, Kenta Oono, Shohei Hido, Justin Clayton, 2015 https://scholar.google.com/scholar?q=Chainer%3A+A+Next-Generation+Open+Source+Framework+for+Deep+Learning 7. Automatic Differentiation in PyTorch — Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, et al., 2017 https://scholar.google.com/scholar?q=Automatic+Differentiation+in+PyTorch 8. Training Deep Nets with Sublinear Memory Cost — Tianqi Chen, Bing Xu, Chiyuan Zhang, Carlos Guestrin, 2016 https://scholar.google.com/scholar?q=Training+Deep+Nets+with+Sublinear+Memory+Cost 9. Reducing Activation Recomputation in Large Transformer Models — Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, Bryan Catanzaro, 2022 https://scholar.google.com/scholar?q=Reducing+Activation+Recomputation+in+Large+Transformer+Models 10. Efficient Rematerialization for Deep Networks — Ravi Kumar, Manish Purohit, Zoya Svitkina, Erik Vee, Joshua Wang, 2019 https://scholar.google.com/scholar?q=Efficient+Rematerialization+for+Deep+Networks 11. Ratel: Optimizing Holistic Data Movement to Fine-Tune 100B Model on a Consumer GPU — Changyue Liao, Mo Sun, Zihan Yang, Jun Xie, Kaiqi Chen, Binhang Yuan, Fei Wu, Zeke Wang, 2025 https://scholar.google.com/scholar?q=Ratel%3A+Optimizing+Holistic+Data+Movement+to+Fine-Tune+100B+Model+on+a+Consumer+GPU 12. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 13. PatrickStar: Parallel Training of Pre-trained Models via a Chunk-based Memory Management — Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu, Jie Zhou, Yang You, 2022 https://scholar.google.com/scholar?q=PatrickStar%3A+Parallel+Training+of+Pre-trained+Models+via+a+Chunk-based+Memory+Management 14. GaLore / other optimizer-state-aware memory-efficient training methods — Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, Yuandong Tian, 2024 https://scholar.google.com/scholar?q=GaLore+%2F+other+optimizer-state-aware+memory-efficient+training+methods Interactive Visualization: MegaTrain: Training 120B-Parameter Models on a Single GPU

  2. 2d ago

    Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

    This episode examines when it actually pays to split LLM inference hardware into four specialized pools rather than the now-standard two-way prefill/decode split, based on a paper proposing PDAF (prefill-attention, prefill-FFN, decode-attention, decode-FFN). It traces the reasoning from why agentic workloads — which can hit context sizes of 100,000+ tokens through repeated tool calls — strain hardware differently than chatbot traffic, through the compute-bound nature of prefill versus the memory-bandwidth-bound nature of decode, and why attention and FFN sublayers batch so differently that combining them on one device forces a similar compromise. It covers prior production systems (DistServe, Splitwise, StepFun's Step-3) that motivated these splits, and introduces the authors' HeteroPanacea simulator, validated against a real 8-node NVIDIA B200 cluster, which they use to search for when the reported up to 2.06x throughput gain actually materializes versus when the added complexity isn't worth it. Listeners interested in LLM serving infrastructure will find a grounded, skeptical take on a systems paper that resists overselling its own headline number. Sources: 1. When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference — Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides, 2026 http://arxiv.org/abs/2608.03741 2. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving 3. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, R. Bianchini, 2024 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 4. Step-3 Is Large yet Affordable: Model-System Co-Design for Cost-Effective Decoding — StepFun, 2025 https://scholar.google.com/scholar?q=Step-3+Is+Large+yet+Affordable%3A+Model-System+Co-Design+for+Cost-Effective+Decoding 5. Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot — R. Qin, Z. Li, W. He, J. Cui, F. Ren, M. Zhang, Y. Wu, W. Zheng, X. Xu, 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-Centric+Architecture+for+Serving+LLM+Chatbot 6. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, Y. Lai, B. Lou, J. Nie, C. Xiao, T. Adeniran, P. Forys, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs Interactive Visualization: Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

  3. 2d ago

    KV Cache Management Faces Its First Head-to-Head Test

    This episode examines a comparative study of three KV cache management strategies for LLM inference — vLLM's PagedAttention memory management, H2O's static sparsification, and InfiniGen's dynamic CPU-offload selection — tested side by side on identical hardware for the first time. The standout finding: both H2O and InfiniGen hit out-of-memory errors around 10,000 tokens, less than 10% of the 128K context window modern models claim to support, revealing that many eviction-based approaches can't even survive prefill on long documents. The discussion traces why KV caches exist at all (avoiding quadratic recomputation cost), how Grouped Query Attention reduces steady-state cache size but does nothing for the transient attention-score matrix that must be materialized during prefill to decide what to evict, and why that structural gap explains the paradigms' divergent failure modes. Testing spans Llama-3.1-8B and 70B, GPT-OSS-20B, and multiple benchmark datasets across four H100 GPUs. Listeners interested in the practical limits of long-context LLM serving — and why architectural tricks like GQA don't fully solve the memory problem — will find the paper's empirical exposure of these failure points compelling. Sources: 1. Comparative Characterization of KV Cache Management Strategies for LLM Inference — Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu, 2026 http://arxiv.org/abs/2604.05012 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — W. Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Z. Zhang et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — W. Lee, J. Lee, J. Seo, J. Sim, 2024 https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management 5. Characterizing the Behavior and Impact of KV Caching on Transformer Inferences under Concurrency — J. Ye, J. Cernuda, A. Maurya, X.-H. Sun, A. Kougkas, B. Nicolae, 2025 https://scholar.google.com/scholar?q=Characterizing+the+Behavior+and+Impact+of+KV+Caching+on+Transformer+Inferences+under+Concurrency 6. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — G. Xiao et al., 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang et al., 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — C. Hooper et al., 2024 https://scholar.google.com/scholar?q=KVQuant%3A+Towards+10+Million+Context+Length+LLM+Inference+with+KV+Cache+Quantization Interactive Visualization: KV Cache Management Faces Its First Head-to-Head Test

  4. 2d ago

    Language Models Learn to Fetch Only What They Need

    This episode examines "Language Models Can Control Their Own Attention" from KAIST AI, which tackles the memory bottleneck of long-context inference: at a million tokens, generating each token requires hauling roughly 15 gigabytes of key-value cache through memory, comparable to reloading the entire model's active parameters. The discussion traces prior fixes—StreamingLLM's attention-sink heuristic, H2O's cumulative attention scoring, and Quest's query-aware page selection—all grouped as "extrinsic scoring" methods that still require a full pass over context statistics before discarding anything. The paper's proposed alternative, Declarative Attention, repurposes chain-of-thought so the model states in its own reasoning which parts of the context it needs, letting the inference engine skip loading the rest instead of relying on an external scorer. The hosts debate whether self-reported relevance is trustworthy compared to an independent extrinsic estimate, since errors here directly create blind spots in what the model can see rather than showing up as recoverable noise. Listeners interested in long-context efficiency, KV-cache management, or the mechanics behind sparse attention will find the back-and-forth over whether this approach is elegant or quietly risky especially engaging. Sources: 1. Language Models Can Control Their Own Attention — Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos, 2026 http://arxiv.org/abs/2609.02737 2. Generating Long Sequences with Sparse Transformers — Rewon Child, Scott Gray, Alec Radford, Ilya Sutskever, 2019 https://scholar.google.com/scholar?q=Generating+Long+Sequences+with+Sparse+Transformers 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023/2024 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 5. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han, 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 6. Self-selected attention span for accelerating large language model inference — Tian Jin, Wanzin Yazar, Zifei Xu, Sayeh Sharify, Xin Wang, 2024 https://scholar.google.com/scholar?q=Self-selected+attention+span+for+accelerating+large+language+model+inference 7. System 2 Attention (is something you might need too) — Jason Weston, Sainbayar Sukhbaatar, 2023 https://scholar.google.com/scholar?q=System+2+Attention+%28is+something+you+might+need+too%29 8. Native sparse attention: Hardware-aligned and natively trainable sparse attention — Jingyang Yuan et al., 2025 https://scholar.google.com/scholar?q=Native+sparse+attention%3A+Hardware-aligned+and+natively+trainable+sparse+attention Interactive Visualization: Language Models Learn to Fetch Only What They Need

  5. 3d ago

    TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

    This episode examines TimesFM, Google Research's decoder-only foundation model for time-series forecasting, and its central claim that a single pretrained 200-million-parameter model can forecast unfamiliar datasets zero-shot, without fine-tuning, at accuracy close to models trained specifically on each dataset. The hosts trace the architecture's lineage from patching, borrowed from the Vision Transformer's image-patch approach and specifically from PatchTST's time-series adaptation, to TimesFM's own contribution of pairing patched inputs with a causal, GPT-style autoregressive setup that naturally handles variable context lengths. They contrast this against DeepAR's RNN-based forecasting, which still required target series in training, and against a 2023 NeurIPS trick of feeding raw numbers as text into large language models, which TimesFM claims to beat at a fraction of the cost. A key surprise is the training corpus itself: since real time-series data is far scarcer online than text, the roughly 100 billion timepoints come largely from Google Trends and Wikipedia pageviews, supplemented by synthetic ARMA and seasonal processes engineered to fill coverage gaps. Listeners interested in foundation models, forecasting infrastructure, or how architectural ideas transfer across modalities will find the discussion's skepticism about benchmark claims and evaluation rigor especially engaging as the hosts preview a closer look at the paper's actual scoring methodology. Sources: 1. TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting https://arxiv.org/pdf/2310.10688 2. https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ 3. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks — David Salinas, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, 2017 (arXiv), 2020 (Intl. J. Forecasting) https://scholar.google.com/scholar?q=DeepAR%3A+Probabilistic+Forecasting+with+Autoregressive+Recurrent+Networks 4. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting — Boris Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio, 2019 (arXiv), ICLR 2020 https://scholar.google.com/scholar?q=N-BEATS%3A+Neural+Basis+Expansion+Analysis+for+Interpretable+Time+Series+Forecasting 5. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — Bryan Lim, Sercan Arik, Nicolas Loeff, Tomas Pfister, 2019 (arXiv), 2021 (Intl. J. Forecasting) https://scholar.google.com/scholar?q=Temporal+Fusion+Transformers+for+Interpretable+Multi-horizon+Time+Series+Forecasting 6. Chronos: Learning the Language of Time Series — Abdul Fatir Ansari, Lorenzo Stella, et al. (Amazon), 2024 https://scholar.google.com/scholar?q=Chronos%3A+Learning+the+Language+of+Time+Series 7. Large Language Models Are Zero-Shot Time Series Forecasters — Nate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon Wilson, 2023 (NeurIPS) https://scholar.google.com/scholar?q=Large+Language+Models+Are+Zero-Shot+Time+Series+Forecasters 8. One Fits All: Power General Time Series Analysis by Pretrained LM — Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, Rong Jin, 2023 (NeurIPS) https://scholar.google.com/scholar?q=One+Fits+All%3A+Power+General+Time+Series+Analysis+by+Pretrained+LM 9. Moirai: Unified Training of Universal Time Series Forecasting Transformers — Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doug Arnold (Salesforce), 2024 (ICML) https://scholar.google.com/scholar?q=Moirai%3A+Unified+Training+of+Universal+Time+Series+Forecasting+Transformers 10. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting — Kashif Rasul, Arjun Ashok, Andrew Robert Williams, et al., 2024 https://scholar.google.com/scholar?q=Lag-Llama%3A+Towards+Foundation+Models+for+Probabilistic+Time+Series+Forecasting 11. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. (Google), 2020 https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale 12. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2023 (ICLR) https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-term+Forecasting+with+Transformers 13. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting — Haoyi Zhou, Shanghang Zhang, Jieqi Peng, et al., 2021 (AAAI, Best Paper) https://scholar.google.com/scholar?q=Informer%3A+Beyond+Efficient+Transformer+for+Long+Sequence+Time-Series+Forecasting 14. Masked Autoencoders Are Scalable Vision Learners — Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, 2021 https://scholar.google.com/scholar?q=Masked+Autoencoders+Are+Scalable+Vision+Learners 15. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers (PatchTST) — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2022 https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-Term+Forecasting+with+Transformers+%28PatchTST%29 16. TimeGPT-1 — Azul Garza, Max Mergenthaler-Canseco, 2023 https://scholar.google.com/scholar?q=TimeGPT-1 17. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann et al., 2022 https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29 Interactive Visualization: TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

  6. 3d ago

    TreeWY: Speculative Verification for Gated DeltaNet Hybrids

    This episode examines TreeWY, a proposed method for speculative decoding verification in hybrid language models that mix standard attention with Gated DeltaNet linear-attention layers. The discussion explains why current systems like vLLM and SGLang must snapshot the full recurrent state at every draft position before verification, since GDN's decay-and-overwrite state update can't be partially rolled back — a cost that multiplies across branches and makes wide speculative draft trees prohibitively memory-expensive. It traces the problem to its root, from the memory-bandwidth bottleneck that motivates speculative decoding in the first place to the mathematical mechanics of the gated delta rule that make hybrid-model states lossy and irreversible. The paper's proposed fix reframes the state update as decayed additive attention with a corrected value, hinting at a way to verify an entire draft tree with a single triangular solve rather than exhaustive snapshotting. Listeners interested in LLM inference efficiency will find the episode's central claim striking: a roughly 128x reduction in per-node memory without sacrificing correctness guarantees, potentially unlocking much more aggressive tree-based speculation on hybrid architectures. Sources: 1. TreeWY: Speculative Verification for Gated DeltaNet Hybrids https://arxiv.org/pdf/2608.20961 2. Bole: Efficient Tree Speculation for Hybrid-Attention Language Models — L. Wang et al., 2026 https://scholar.google.com/scholar?q=Bole%3A+Efficient+Tree+Speculation+for+Hybrid-Attention+Language+Models 3. ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab and NVIDIA, 2026 https://scholar.google.com/scholar?q=ReplaySSM%3A+Cache+SSM+Inputs%2C+Not+State 4. STree: Speculative Tree Decoding for Hybrid State-Space Models — Y. Wu et al., 2025 https://scholar.google.com/scholar?q=STree%3A+Speculative+Tree+Decoding+for+Hybrid+State-Space+Models 5. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang et al., 2024 https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length 6. Gated Delta Networks: Improving Mamba2 with Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025 https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule

  7. 4d ago

    Reptile: The First-Order Meta-Learning Shortcut That Works

    This episode examines "On First-Order Meta-Learning Algorithms" by Alex Nichol, Joshua Achiam, and John Schulman, which challenges the assumption that MAML's expensive second-derivative computation is essential for effective few-shot learning. The discussion traces the lineage from MAML's nested optimization — where an outer loop backpropagates through an inner loop's gradient steps via the Hessian — through First-Order MAML's approximation, to Reptile, a stripped-down algorithm that simply runs SGD on sampled tasks and nudges the initialization toward the result, with no meta-gradient or train-test split required. A central tension drives the conversation: why pulling an initialization toward "wherever SGD landed" produces a genuinely different target than plain joint training across tasks, rather than just averaging into one generic model. The hosts set up a Taylor-expansion argument to explain which gradient terms MAML, FOMAML, and Reptile weight differently, revealing the mathematical reason the cheaper approximation retains nearly all the useful signal. Listeners interested in the mechanics of meta-learning, gradient-based optimization tradeoffs, or the history of few-shot learning approaches will find the paper's practical implications for scaling meta-learning algorithms especially relevant. Sources: 1. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018 http://arxiv.org/abs/1803.02999 2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017 https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks 3. How to train your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019 https://scholar.google.com/scholar?q=How+to+train+your+MAML 4. Meta-Learning with Implicit Gradients — Aravind Rajeswaran, Chelsea Finn, Sham Kakade, Sergey Levine, 2019 https://scholar.google.com/scholar?q=Meta-Learning+with+Implicit+Gradients 5. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017 https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning 6. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016 https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent 7. Using Fast Weights to Deblur Old Memories — Geoffrey E. Hinton, David C. Plaut, 1987 https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Deblur+Old+Memories 8. Parallelized Stochastic Gradient Descent — Martin Zinkevich, Markus Weimer, Lihong Li, Alex J. Smola, 2010 https://scholar.google.com/scholar?q=Parallelized+Stochastic+Gradient+Descent

  8. Aug 30

    FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

    This episode examines FLINT, a proposed hardware/software architecture from Huawei's Zurich research lab (with ETH Zürich and HUST) for closing the gap between multi-terabyte LLM weight sizes and the far smaller on-package memory of today's GPUs. It introduces high bandwidth flash (HBF), an emerging memory tier that stacks 3D NAND dies with through-silicon vias to sit directly beside HBM in the accelerator package, storing read-only model weights while HBM handles fast-changing KV cache and activations. The discussion walks through core NAND flash mechanics — dies, planes, blocks, and pages, along with the punishing asymmetry between microsecond reads and millisecond erases — to explain why naive flash designs stall under refresh operations and static prefetching. It then details how FLINT's burst-buffer controller replaces compiler-driven prefetch hints with real-time demand-based read coalescing, using HBF's built-in page and cache buffers instead of dedicated SRAM. Listeners interested in memory system design, inference hardware economics, and the practical engineering trade-offs of scaling capacity without wasting GPU compute will find this a detailed look at a genuinely emerging technology rather than a shipping product. Sources: 1. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration — Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang, 2026 http://arxiv.org/abs/2608.25062 2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System — Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, Jinho Lee, 2024, HPCA https://scholar.google.com/scholar?q=Smart-Infinity%3A+Fast+Large+Language+Model+Training+using+Near-Storage+Processing+on+a+Real+System 5. InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference — Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang, 2024 https://scholar.google.com/scholar?q=InstInfer%3A+In-Storage+Attention+Offloading+for+Cost-Effective+Long-Context+LLM+Inference 6. DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings — Aayush Gupta, Youngjae Kim, Bhuvan Urgaonkar, 2009, ASPLOS https://scholar.google.com/scholar?q=DFTL%3A+A+Flash+Translation+Layer+Employing+Demand-based+Selective+Caching+of+Page-level+Address+Mappings 7. Design Tradeoffs for SSD Performance — Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D. Davis, Mark Manasse, Rina Panigrahy, 2008, USENIX ATC https://scholar.google.com/scholar?q=Design+Tradeoffs+for+SSD+Performance 8. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs — Matias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh, Damien Le Moal, Gregory R. Ganger, George Amvrosiadis, 2021, USENIX ATC https://scholar.google.com/scholar?q=ZNS%3A+Avoiding+the+Block+Interface+Tax+for+Flash-based+SSDs 9. RAIDR: Retention-Aware Intelligent DRAM Refresh — Jamie Liu, Ben Jaiyen, Richard Veras, Onur Mutlu, 2012, ISCA https://scholar.google.com/scholar?q=RAIDR%3A+Retention-Aware+Intelligent+DRAM+Refresh 10. Threshold Voltage Distribution in MLC NAND Flash Memory: Characterization, Analysis, and Modeling — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, DATE https://scholar.google.com/scholar?q=Threshold+Voltage+Distribution+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Analysis%2C+and+Modeling 11. Error Analysis and Retention-Aware Error Management for NAND Flash Memory — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, Intel Technology Journal https://scholar.google.com/scholar?q=Error+Analysis+and+Retention-Aware+Error+Management+for+NAND+Flash+Memory 12. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — M. Ha, E. Kim, H. Kim, 2026 https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference 13. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — K. Alizadeh, S. I. Mirzadeh, et al., 2024 https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory 14. PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving — A. C. Yüzügüler, J. Zhuang, L. Cavigelli, 2025 https://scholar.google.com/scholar?q=PRESERVE%3A+Prefetching+Model+Weights+and+KV-Cache+in+Distributed+LLM+Serving 15. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs 16. Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory — W. Sun, M. Gao, et al., 2025 https://scholar.google.com/scholar?q=Lincoln%3A+Real-Time+50~100B+LLM+Inference+on+Consumer+Devices+with+LPDDR-Interfaced%2C+Compute-Enabled+Flash+Memory Interactive Visualization: FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

Ratings & Reviews

3.7
out of 5
3 Ratings

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like