AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 23 giờ trước

    Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

    This episode examines when it actually pays to split LLM inference hardware into four specialized pools rather than the now-standard two-way prefill/decode split, based on a paper proposing PDAF (prefill-attention, prefill-FFN, decode-attention, decode-FFN). It traces the reasoning from why agentic workloads — which can hit context sizes of 100,000+ tokens through repeated tool calls — strain hardware differently than chatbot traffic, through the compute-bound nature of prefill versus the memory-bandwidth-bound nature of decode, and why attention and FFN sublayers batch so differently that combining them on one device forces a similar compromise. It covers prior production systems (DistServe, Splitwise, StepFun's Step-3) that motivated these splits, and introduces the authors' HeteroPanacea simulator, validated against a real 8-node NVIDIA B200 cluster, which they use to search for when the reported up to 2.06x throughput gain actually materializes versus when the added complexity isn't worth it. Listeners interested in LLM serving infrastructure will find a grounded, skeptical take on a systems paper that resists overselling its own headline number. Sources: 1. When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference — Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides, 2026 http://arxiv.org/abs/2608.03741 2. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving 3. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, R. Bianchini, 2024 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 4. Step-3 Is Large yet Affordable: Model-System Co-Design for Cost-Effective Decoding — StepFun, 2025 https://scholar.google.com/scholar?q=Step-3+Is+Large+yet+Affordable%3A+Model-System+Co-Design+for+Cost-Effective+Decoding 5. Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot — R. Qin, Z. Li, W. He, J. Cui, F. Ren, M. Zhang, Y. Wu, W. Zheng, X. Xu, 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-Centric+Architecture+for+Serving+LLM+Chatbot 6. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, Y. Lai, B. Lou, J. Nie, C. Xiao, T. Adeniran, P. Forys, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs Interactive Visualization: Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

  2. 23 giờ trước

    KV Cache Management Faces Its First Head-to-Head Test

    This episode examines a comparative study of three KV cache management strategies for LLM inference — vLLM's PagedAttention memory management, H2O's static sparsification, and InfiniGen's dynamic CPU-offload selection — tested side by side on identical hardware for the first time. The standout finding: both H2O and InfiniGen hit out-of-memory errors around 10,000 tokens, less than 10% of the 128K context window modern models claim to support, revealing that many eviction-based approaches can't even survive prefill on long documents. The discussion traces why KV caches exist at all (avoiding quadratic recomputation cost), how Grouped Query Attention reduces steady-state cache size but does nothing for the transient attention-score matrix that must be materialized during prefill to decide what to evict, and why that structural gap explains the paradigms' divergent failure modes. Testing spans Llama-3.1-8B and 70B, GPT-OSS-20B, and multiple benchmark datasets across four H100 GPUs. Listeners interested in the practical limits of long-context LLM serving — and why architectural tricks like GQA don't fully solve the memory problem — will find the paper's empirical exposure of these failure points compelling. Sources: 1. Comparative Characterization of KV Cache Management Strategies for LLM Inference — Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu, 2026 http://arxiv.org/abs/2604.05012 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — W. Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Z. Zhang et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — W. Lee, J. Lee, J. Seo, J. Sim, 2024 https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management 5. Characterizing the Behavior and Impact of KV Caching on Transformer Inferences under Concurrency — J. Ye, J. Cernuda, A. Maurya, X.-H. Sun, A. Kougkas, B. Nicolae, 2025 https://scholar.google.com/scholar?q=Characterizing+the+Behavior+and+Impact+of+KV+Caching+on+Transformer+Inferences+under+Concurrency 6. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — G. Xiao et al., 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang et al., 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — C. Hooper et al., 2024 https://scholar.google.com/scholar?q=KVQuant%3A+Towards+10+Million+Context+Length+LLM+Inference+with+KV+Cache+Quantization Interactive Visualization: KV Cache Management Faces Its First Head-to-Head Test

  3. 1 ngày trước

    TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

    This episode examines TimesFM, Google Research's decoder-only foundation model for time-series forecasting, and its central claim that a single pretrained 200-million-parameter model can forecast unfamiliar datasets zero-shot, without fine-tuning, at accuracy close to models trained specifically on each dataset. The hosts trace the architecture's lineage from patching, borrowed from the Vision Transformer's image-patch approach and specifically from PatchTST's time-series adaptation, to TimesFM's own contribution of pairing patched inputs with a causal, GPT-style autoregressive setup that naturally handles variable context lengths. They contrast this against DeepAR's RNN-based forecasting, which still required target series in training, and against a 2023 NeurIPS trick of feeding raw numbers as text into large language models, which TimesFM claims to beat at a fraction of the cost. A key surprise is the training corpus itself: since real time-series data is far scarcer online than text, the roughly 100 billion timepoints come largely from Google Trends and Wikipedia pageviews, supplemented by synthetic ARMA and seasonal processes engineered to fill coverage gaps. Listeners interested in foundation models, forecasting infrastructure, or how architectural ideas transfer across modalities will find the discussion's skepticism about benchmark claims and evaluation rigor especially engaging as the hosts preview a closer look at the paper's actual scoring methodology. Sources: 1. TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting https://arxiv.org/pdf/2310.10688 2. https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ 3. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks — David Salinas, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, 2017 (arXiv), 2020 (Intl. J. Forecasting) https://scholar.google.com/scholar?q=DeepAR%3A+Probabilistic+Forecasting+with+Autoregressive+Recurrent+Networks 4. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting — Boris Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio, 2019 (arXiv), ICLR 2020 https://scholar.google.com/scholar?q=N-BEATS%3A+Neural+Basis+Expansion+Analysis+for+Interpretable+Time+Series+Forecasting 5. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — Bryan Lim, Sercan Arik, Nicolas Loeff, Tomas Pfister, 2019 (arXiv), 2021 (Intl. J. Forecasting) https://scholar.google.com/scholar?q=Temporal+Fusion+Transformers+for+Interpretable+Multi-horizon+Time+Series+Forecasting 6. Chronos: Learning the Language of Time Series — Abdul Fatir Ansari, Lorenzo Stella, et al. (Amazon), 2024 https://scholar.google.com/scholar?q=Chronos%3A+Learning+the+Language+of+Time+Series 7. Large Language Models Are Zero-Shot Time Series Forecasters — Nate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon Wilson, 2023 (NeurIPS) https://scholar.google.com/scholar?q=Large+Language+Models+Are+Zero-Shot+Time+Series+Forecasters 8. One Fits All: Power General Time Series Analysis by Pretrained LM — Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, Rong Jin, 2023 (NeurIPS) https://scholar.google.com/scholar?q=One+Fits+All%3A+Power+General+Time+Series+Analysis+by+Pretrained+LM 9. Moirai: Unified Training of Universal Time Series Forecasting Transformers — Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doug Arnold (Salesforce), 2024 (ICML) https://scholar.google.com/scholar?q=Moirai%3A+Unified+Training+of+Universal+Time+Series+Forecasting+Transformers 10. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting — Kashif Rasul, Arjun Ashok, Andrew Robert Williams, et al., 2024 https://scholar.google.com/scholar?q=Lag-Llama%3A+Towards+Foundation+Models+for+Probabilistic+Time+Series+Forecasting 11. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. (Google), 2020 https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale 12. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2023 (ICLR) https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-term+Forecasting+with+Transformers 13. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting — Haoyi Zhou, Shanghang Zhang, Jieqi Peng, et al., 2021 (AAAI, Best Paper) https://scholar.google.com/scholar?q=Informer%3A+Beyond+Efficient+Transformer+for+Long+Sequence+Time-Series+Forecasting 14. Masked Autoencoders Are Scalable Vision Learners — Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, 2021 https://scholar.google.com/scholar?q=Masked+Autoencoders+Are+Scalable+Vision+Learners 15. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers (PatchTST) — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2022 https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-Term+Forecasting+with+Transformers+%28PatchTST%29 16. TimeGPT-1 — Azul Garza, Max Mergenthaler-Canseco, 2023 https://scholar.google.com/scholar?q=TimeGPT-1 17. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann et al., 2022 https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29 Interactive Visualization: TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

  4. 1 ngày trước

    TreeWY: Speculative Verification for Gated DeltaNet Hybrids

    This episode examines TreeWY, a proposed method for speculative decoding verification in hybrid language models that mix standard attention with Gated DeltaNet linear-attention layers. The discussion explains why current systems like vLLM and SGLang must snapshot the full recurrent state at every draft position before verification, since GDN's decay-and-overwrite state update can't be partially rolled back — a cost that multiplies across branches and makes wide speculative draft trees prohibitively memory-expensive. It traces the problem to its root, from the memory-bandwidth bottleneck that motivates speculative decoding in the first place to the mathematical mechanics of the gated delta rule that make hybrid-model states lossy and irreversible. The paper's proposed fix reframes the state update as decayed additive attention with a corrected value, hinting at a way to verify an entire draft tree with a single triangular solve rather than exhaustive snapshotting. Listeners interested in LLM inference efficiency will find the episode's central claim striking: a roughly 128x reduction in per-node memory without sacrificing correctness guarantees, potentially unlocking much more aggressive tree-based speculation on hybrid architectures. Sources: 1. TreeWY: Speculative Verification for Gated DeltaNet Hybrids https://arxiv.org/pdf/2608.20961 2. Bole: Efficient Tree Speculation for Hybrid-Attention Language Models — L. Wang et al., 2026 https://scholar.google.com/scholar?q=Bole%3A+Efficient+Tree+Speculation+for+Hybrid-Attention+Language+Models 3. ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab and NVIDIA, 2026 https://scholar.google.com/scholar?q=ReplaySSM%3A+Cache+SSM+Inputs%2C+Not+State 4. STree: Speculative Tree Decoding for Hybrid State-Space Models — Y. Wu et al., 2025 https://scholar.google.com/scholar?q=STree%3A+Speculative+Tree+Decoding+for+Hybrid+State-Space+Models 5. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang et al., 2024 https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length 6. Gated Delta Networks: Improving Mamba2 with Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025 https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule

  5. 2 ngày trước

    Reptile: The First-Order Meta-Learning Shortcut That Works

    This episode examines "On First-Order Meta-Learning Algorithms" by Alex Nichol, Joshua Achiam, and John Schulman, which challenges the assumption that MAML's expensive second-derivative computation is essential for effective few-shot learning. The discussion traces the lineage from MAML's nested optimization — where an outer loop backpropagates through an inner loop's gradient steps via the Hessian — through First-Order MAML's approximation, to Reptile, a stripped-down algorithm that simply runs SGD on sampled tasks and nudges the initialization toward the result, with no meta-gradient or train-test split required. A central tension drives the conversation: why pulling an initialization toward "wherever SGD landed" produces a genuinely different target than plain joint training across tasks, rather than just averaging into one generic model. The hosts set up a Taylor-expansion argument to explain which gradient terms MAML, FOMAML, and Reptile weight differently, revealing the mathematical reason the cheaper approximation retains nearly all the useful signal. Listeners interested in the mechanics of meta-learning, gradient-based optimization tradeoffs, or the history of few-shot learning approaches will find the paper's practical implications for scaling meta-learning algorithms especially relevant. Sources: 1. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018 http://arxiv.org/abs/1803.02999 2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017 https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks 3. How to train your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019 https://scholar.google.com/scholar?q=How+to+train+your+MAML 4. Meta-Learning with Implicit Gradients — Aravind Rajeswaran, Chelsea Finn, Sham Kakade, Sergey Levine, 2019 https://scholar.google.com/scholar?q=Meta-Learning+with+Implicit+Gradients 5. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017 https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning 6. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016 https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent 7. Using Fast Weights to Deblur Old Memories — Geoffrey E. Hinton, David C. Plaut, 1987 https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Deblur+Old+Memories 8. Parallelized Stochastic Gradient Descent — Martin Zinkevich, Markus Weimer, Lihong Li, Alex J. Smola, 2010 https://scholar.google.com/scholar?q=Parallelized+Stochastic+Gradient+Descent

  6. 5 ngày trước

    FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

    This episode examines FLINT, a proposed hardware/software architecture from Huawei's Zurich research lab (with ETH Zürich and HUST) for closing the gap between multi-terabyte LLM weight sizes and the far smaller on-package memory of today's GPUs. It introduces high bandwidth flash (HBF), an emerging memory tier that stacks 3D NAND dies with through-silicon vias to sit directly beside HBM in the accelerator package, storing read-only model weights while HBM handles fast-changing KV cache and activations. The discussion walks through core NAND flash mechanics — dies, planes, blocks, and pages, along with the punishing asymmetry between microsecond reads and millisecond erases — to explain why naive flash designs stall under refresh operations and static prefetching. It then details how FLINT's burst-buffer controller replaces compiler-driven prefetch hints with real-time demand-based read coalescing, using HBF's built-in page and cache buffers instead of dedicated SRAM. Listeners interested in memory system design, inference hardware economics, and the practical engineering trade-offs of scaling capacity without wasting GPU compute will find this a detailed look at a genuinely emerging technology rather than a shipping product. Sources: 1. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration — Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang, 2026 http://arxiv.org/abs/2608.25062 2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System — Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, Jinho Lee, 2024, HPCA https://scholar.google.com/scholar?q=Smart-Infinity%3A+Fast+Large+Language+Model+Training+using+Near-Storage+Processing+on+a+Real+System 5. InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference — Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang, 2024 https://scholar.google.com/scholar?q=InstInfer%3A+In-Storage+Attention+Offloading+for+Cost-Effective+Long-Context+LLM+Inference 6. DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings — Aayush Gupta, Youngjae Kim, Bhuvan Urgaonkar, 2009, ASPLOS https://scholar.google.com/scholar?q=DFTL%3A+A+Flash+Translation+Layer+Employing+Demand-based+Selective+Caching+of+Page-level+Address+Mappings 7. Design Tradeoffs for SSD Performance — Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D. Davis, Mark Manasse, Rina Panigrahy, 2008, USENIX ATC https://scholar.google.com/scholar?q=Design+Tradeoffs+for+SSD+Performance 8. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs — Matias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh, Damien Le Moal, Gregory R. Ganger, George Amvrosiadis, 2021, USENIX ATC https://scholar.google.com/scholar?q=ZNS%3A+Avoiding+the+Block+Interface+Tax+for+Flash-based+SSDs 9. RAIDR: Retention-Aware Intelligent DRAM Refresh — Jamie Liu, Ben Jaiyen, Richard Veras, Onur Mutlu, 2012, ISCA https://scholar.google.com/scholar?q=RAIDR%3A+Retention-Aware+Intelligent+DRAM+Refresh 10. Threshold Voltage Distribution in MLC NAND Flash Memory: Characterization, Analysis, and Modeling — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, DATE https://scholar.google.com/scholar?q=Threshold+Voltage+Distribution+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Analysis%2C+and+Modeling 11. Error Analysis and Retention-Aware Error Management for NAND Flash Memory — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, Intel Technology Journal https://scholar.google.com/scholar?q=Error+Analysis+and+Retention-Aware+Error+Management+for+NAND+Flash+Memory 12. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — M. Ha, E. Kim, H. Kim, 2026 https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference 13. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — K. Alizadeh, S. I. Mirzadeh, et al., 2024 https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory 14. PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving — A. C. Yüzügüler, J. Zhuang, L. Cavigelli, 2025 https://scholar.google.com/scholar?q=PRESERVE%3A+Prefetching+Model+Weights+and+KV-Cache+in+Distributed+LLM+Serving 15. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs 16. Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory — W. Sun, M. Gao, et al., 2025 https://scholar.google.com/scholar?q=Lincoln%3A+Real-Time+50~100B+LLM+Inference+on+Consumer+Devices+with+LPDDR-Interfaced%2C+Compute-Enabled+Flash+Memory Interactive Visualization: FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

  7. 5 ngày trước

    LinearKV: When Exact State Merging Breaks Hybrid Models

    This episode examines LinearKV, a new approach to position-independent caching for hybrid Mamba-attention language models, and a striking result: the mathematically "correct" way to merge cached context — exact algebraic composition of recurrent states — badly degrades one tested model's output quality, while a simpler shortcut using only the most recent cached chunk performs reliably. The discussion contrasts standard prefix caching (used by vLLM's PagedAttention and SGLang's RadixAttention) with position-independent caching, which lets cached chunks be reused regardless of order, and explains why that trick breaks down for linear-recurrent layers like Mamba-2 and Gated DeltaNet, which compress history into a single fixed-size state rather than a token-indexed KV ledger. It traces the core problem to how each cached chunk's recurrent state was built in isolation, making "exact" composition exact only relative to a flawed reference rather than the true full-context computation. Listeners interested in LLM serving infrastructure, caching systems, or the tradeoffs of hybrid architectures will find this a concrete case study in how systems intuitions from full attention can actively mislead when applied to newer recurrent designs. Sources: 1. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs — Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li, 2026 http://arxiv.org/abs/2608.11231 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024 https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs 4. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2025 https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion 5. EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models — Junhao Hu et al., 2025 https://scholar.google.com/scholar?q=EPIC%3A+Efficient+Position-Independent+Context+Caching+for+Serving+Large+Language+Models 6. HYPIC: Accelerating hybrid-attention LLM serving with position-independent caching — Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen, 2026 https://scholar.google.com/scholar?q=HYPIC%3A+Accelerating+hybrid-attention+LLM+serving+with+position-independent+caching 7. Marconi: Prefix caching for the era of hybrid LLMs — Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, Ravi Netravali, 2025 https://scholar.google.com/scholar?q=Marconi%3A+Prefix+caching+for+the+era+of+hybrid+LLMs 8. Gated Delta Networks: Improving Mamba2 with the Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2025 (ICLR) https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+the+Delta+Rule 9. EPIC: Efficient position-independent caching for serving large language models — Junhao Hu, Wenrui Huang, Weidong Wang, et al., 2025 (ICML) https://scholar.google.com/scholar?q=EPIC%3A+Efficient+position-independent+caching+for+serving+large+language+models Interactive Visualization: LinearKV: When Exact State Merging Breaks Hybrid Models

  8. 28 thg 8

    Decoupling KL Direction from Rollout Source in LLM Distillation

    This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question. Sources: 1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026 http://arxiv.org/abs/2605.16826 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015 https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks 4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016 https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks 5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026 https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models 8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025 https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation

Xếp Hạng & Nhận Xét

3,7
/5
3 Xếp hạng

Giới Thiệu

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

Có Thể Bạn Cũng Thích