AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. −2 d

    TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

    This episode explores TFGN, an architectural approach to continual pre-training of large language models that claims to solve catastrophic forgetting without four common crutches: replay buffers, task identifiers, small-scale toy benchmarks, and external penalty terms like Fisher-information regularization. The hosts trace the lineage of the forgetting problem back to 1989, explain why popular fixes like LoRA-based parameter-efficient fine-tuning don't actually address forgetting (they just shrink the blast radius), and why classic regularization methods like Elastic Weight Consolidation break down at billion-parameter scale. They also clarify why long-context windows and prompt-based knowledge aren't a substitute for genuinely updating model weights on massive, unbounded corpora like full codebases or legal archives. The conversation lays out TFGN's core mechanism as a dense, input-conditioned overlay operating inside each transformer block, contrasting it with sparse mixture-of-experts routing, and sets up backward transfer as the key metric for measuring whether old knowledge survives new training. Listeners interested in how production LLMs might eventually absorb new domains without expensive retraining or fragile adapter stacking will find the framing of this open problem sharply drawn. Sources: 1. TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale — Anurup Ganguli, 2026 http://arxiv.org/abs/2605.15053 2. Overcoming catastrophic forgetting in neural networks (EWC) — J. Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+catastrophic+forgetting+in+neural+networks+%28EWC%29 3. Loss of plasticity in deep continual learning — S. Dohare et al., 2024, Nature https://scholar.google.com/scholar?q=Loss+of+plasticity+in+deep+continual+learning 4. Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning — O. Y. L. Imanov, 2026, arXiv:2601.18699 https://scholar.google.com/scholar?q=Mechanistic+Analysis+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-Tuning 5. Examining Forgetting in Continual Pre-training of Aligned Large Language Models — C.-A. Li and H.-Y. Lee, 2024, arXiv:2401.03129 https://scholar.google.com/scholar?q=Examining+Forgetting+in+Continual+Pre-training+of+Aligned+Large+Language+Models 6. Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models — I. Abbes, G. Subbaraj, M. Riemer, et al., 2025, arXiv:2508.01908 https://scholar.google.com/scholar?q=Revisiting+Replay+and+Gradient+Alignment+for+Continual+Pre-Training+of+Large+Language+Models 7. Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought — Y. Zhang, B. Tang, T. Ju, S. Duan, G. Liu, 2025, arXiv:2512.21711 https://scholar.google.com/scholar?q=Do+Latent+Tokens+Think%3F+A+Causal+and+Adversarial+Analysis+of+Chain-of-Continuous-Thought Interactive Visualization: TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

  2. −3 d

    SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

    This episode explores SnapStream, a technique from SambaNova Systems for compressing KV caches during long-sequence LLM decoding on dataflow accelerators, demonstrated at production scale with a 671-billion-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens per second. The discussion covers why established training-free KV cache eviction methods like SnapKV and StreamingLLM have struggled to reach real deployments despite promising accuracy results: continuous batching makes it unclear when to trigger compression across requests at different lifecycle stages, and static-graph compilers used by dataflow accelerators can't easily accommodate the dynamic, variable-shaped operations that standard compression implementations rely on. It explains how SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window approach into a single fixed-size cache, splitting sequences into sink tokens, recent tokens, and a compressed middle section during prefill. The conversation is grounded in fundamentals—clarifying the prefill/decode split, why decode is memory-bound, and what makes dataflow accelerators architecturally different from GPUs—making it accessible to listeners unfamiliar with KV cache mechanics while still delivering a specific, hardware-grounded engineering story rather than a purely algorithmic one. Sources: 1. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, Evgenii Iuliugin, Magnus Vesterlund, Christian Häggström, Guangtao Wang, Shubhangi Upasani, Ayush Sachdeva, Rui Li, Faline Fu, Chen Wu, Ayesha Siddiqua, John Long, Tuowen Zhao, Matheen Musaddiq, Håkan Zeffer, Yun Du, Mingran Wang, Qinghua Li, Bo Li, Urmish Thakker, Raghu Prabhakar, 2025 http://arxiv.org/abs/2511.03092 2. Plasticine: A Reconfigurable Architecture For Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+For+Parallel+Patterns 3. SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — Raghu Prabhakar and SambaNova Systems architecture team, 2024 https://scholar.google.com/scholar?q=SN40L%3A+Scaling+the+AI+Memory+Wall+with+Dataflow+and+Composition+of+Experts 4. Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, et al. (Groq), 2020 https://scholar.google.com/scholar?q=Think+Fast%3A+A+Tensor+Streaming+Processor+%28TSP%29+for+Accelerating+Deep+Learning+Workloads 5. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 6. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, S. Han, 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, M. Sun, 2024 https://scholar.google.com/scholar?q=InfLLM%3A+Training-Free+Long-Context+Extrapolation+with+an+Efficient+Context+Memory 9. DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025 https://scholar.google.com/scholar?q=DeepSeek-V3.2-Exp%3A+Boosting+Long-Context+Efficiency+with+DeepSeek+Sparse+Attention 10. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha, D. Zinsley, E. Liu, W. Tennien, A. Rudra, J. Zou, A. Mirhoseini, C. Re, 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study 11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — G. Wang, S. Upasani, C. Wu, D. Gandhi, J. Li, C. Hu, B. Li, U. Thakker, 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+%28SAGE-KV%29 Interactive Visualization: SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

  3. −3 d

    StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

    This episode explores StrataCL, a fabric-native communication library from researchers at Peking University, ICT-CAS, UCAS, Shanghai Jiao Tong University, and Huawei, tested on Huawei's CloudMatrix384 supernode. The discussion centers on how communication overhead — which the paper puts at 30-45% of end-to-end time in distributed LLM training and up to 50% at scale — can be cut by giving collectives and MoE routing true zero-copy access to application buffers on unified-memory fabrics, without breaking compatibility with frameworks like PyTorch and SGLang. A key insight is why buffer-centric libraries like NCCL and HCCL fall short even on fast unified-address fabrics, and how MoE dispatch/combine traffic exposes the limits of naive redesigns. The core technical contribution is registration-on-allocation: exploiting the multi-second gap between physical memory allocation and first use by a communication operator to move registration off the critical path entirely, asynchronously, the moment memory is mapped. The result is a 1.4x iteration-time speedup on a 512-die production training run with no changes to the model, optimizer, or data — pure systems engineering payoff. Sources: 1. StrataCL: Fabric-Native Communication Library for Production Supernodes — Tiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang, 2026 http://arxiv.org/abs/2607.26444 2. U-Net: A User-Level Network Interface for Parallel and Distributed Computing — Thorsten von Eicken, Anindya Basu, Vineet Buch, Werner Vogels, 1995 https://scholar.google.com/scholar?q=U-Net%3A+A+User-Level+Network+Interface+for+Parallel+and+Distributed+Computing 3. Design Guidelines for High Performance RDMA Systems — Anuj Kalia, Michael Kaminsky, David G. Andersen, 2016 https://scholar.google.com/scholar?q=Design+Guidelines+for+High+Performance+RDMA+Systems 4. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, Ion Stoica, 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML 5. NVSHMEM (GPU-initiated, PGAS-style one-sided communication library) — NVIDIA (library/runtime, not a single academic paper), 2016 (initial release, iterated since) https://scholar.google.com/scholar?q=NVSHMEM+%28GPU-initiated%2C+PGAS-style+one-sided+communication+library%29 6. Collective Communication for 100k+ GPUs — Min Si, Pavan Balaji, Yongzhou Chen, et al., 2025 https://scholar.google.com/scholar?q=Collective+Communication+for+100k%2B+GPUs 7. SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading — Xingyi Li, Yadong Liu, Xiaojie Huang, et al., 2026 (NSDI 26) https://scholar.google.com/scholar?q=SwiftEP%3A+Accelerating+MoE+Inference+with+Buffer+Fusion+and+TMA+Offloading 8. PyTorch Symmetric Memory / NVSHMEM-style same-VA mirrored buffers — PyTorch Team / NVIDIA (NVSHMEM), 2024-2025 https://scholar.google.com/scholar?q=PyTorch+Symmetric+Memory+%2F+NVSHMEM-style+same-VA+mirrored+buffers 9. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 (SOSP 23) https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention Interactive Visualization: StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

  4. −4 d

    MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

    This episode explores a challenge to conventional wisdom in parameter-efficient fine-tuning, examining a method called MiCA that inverts the logic behind LoRA (Low-Rank Adaptation). Rather than letting trainable weight-update matrices drift freely, as standard LoRA does, MiCA deliberately anchors one matrix to the minor singular-value directions of a weight matrix — the low-energy, rarely-used "corners" that classical compression theory says to discard — leaving those directions free for new knowledge rather than overwriting the dominant, pretrained-heavy subspace. The discussion traces the technique's lineage through SVD, the Eckart-Young-Mirsky theorem, P***A's SVD-based initialization, and Minor Component Analysis, framing MiCA's core bet: catastrophic forgetting during fine-tuning may stem from cramming new information into already-saturated high-energy directions. Listeners interested in the mechanics of efficient model adaptation, knowledge editing, and where the field's assumptions about "useless" weight-matrix structure might be wrong will find the debate over whether this is a genuine architectural insight or a narrower refinement of existing PEFT ideas especially engaging. Sources: 1. MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA https://arxiv.org/pdf/2604.01694 2. The Approximation of One Matrix by Another of Lower Rank — Carl Eckart, Gale Young, 1936 https://scholar.google.com/scholar?q=The+Approximation+of+One+Matrix+by+Another+of+Lower+Rank 3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models 4. P***A: Principal Singular Values and Singular Vectors Adaptation of Large Language Models — Fanxu Meng, Zhaohui Wang, Muhan Zhang, 2024 https://scholar.google.com/scholar?q=PiSSA%3A+Principal+Singular+Values+and+Singular+Vectors+Adaptation+of+Large+Language+Models 5. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning — Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, Tuo Zhao, 2023 https://scholar.google.com/scholar?q=AdaLoRA%3A+Adaptive+Budget+Allocation+for+Parameter-Efficient+Fine-Tuning 6. SOMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning — Seokju Yun, Seunghye Chae, Dongheon Lee, Youngmin Ro, 2025 https://scholar.google.com/scholar?q=SOMA%3A+Singular+Value+Decomposed+Minor+Components+Adaptation+for+Domain+Generalizable+Representation+Learning 7. Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-Tuning — Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh, 2026 https://scholar.google.com/scholar?q=Learning+Rate+Matters%3A+Vanilla+LoRA+May+Suffice+for+LLM+Fine-Tuning 8. DoRA: Weight-Decomposed Low-Rank Adaptation — Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen, 2024 https://scholar.google.com/scholar?q=DoRA%3A+Weight-Decomposed+Low-Rank+Adaptation 9. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 10. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic Interactive Visualization: MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

  5. −4 d

    Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing

    This episode explores cross-instance attention in disaggregated LLM serving, focusing on the surprising size inversion created by Multi-head Latent Attention: a routed decoding query shrinks to roughly a kilobyte while the cache chunk it must read can balloon to 61 megabytes across layers, upending the old assumption that query and cache are comparably sized. The discussion traces why this scenario is becoming routine — providers sharing precomputed caches for large corpora that outgrow a single GPU's memory, and agentic workloads where many sub-agents query one oversized shared prefix — and lays out the three possible strategies (route, fetch, or recompute locally) for handling the mismatch, including how sparse indexers further shrink the routable unit to scattered top-k blocks. A key thread examines device-initiated RDMA via IBGDA, challenging the intuition that skipping the CPU proxy is automatically faster: prior work on tiny mixture-of-experts messages actually found IBGDA slower, but the paper's controlled test on kilobyte-scale attention traffic shows the CPU-proxy path is 40% slower at the median and over 50% slower at steady state. Listeners interested in GPU networking, KV-cache architecture, or the practical plumbing behind large-scale LLM inference will find the paper's empirical resolution of a previously untested assumption particularly compelling. Sources: 1. Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics — Bole Ma, Jan Eitzinger, Harald Köstler, Gerhard Wellein, 2026 http://arxiv.org/abs/2606.01502 2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 3. DeepSeek-V3 Technical Report — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023 https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 5. TransMLA: Multi-Head Latent Attention Is All You Need — Fanxu Meng, Zengwei Yao, Muhan Zhang, et al., 2025 https://scholar.google.com/scholar?q=TransMLA%3A+Multi-Head+Latent+Attention+Is+All+You+Need 6. Improving Network Performance of HPC Systems Using NVIDIA Magnum IO NVSHMEM and GPUDirect Async — NVIDIA (NVSHMEM / Magnum IO engineering team), 2023 https://scholar.google.com/scholar?q=Improving+Network+Performance+of+HPC+Systems+Using+NVIDIA+Magnum+IO+NVSHMEM+and+GPUDirect+Async 7. Efficient Inter-node MPI Communication using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs — Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda, 2013 https://scholar.google.com/scholar?q=Efficient+Inter-node+MPI+Communication+using+GPUDirect+RDMA+for+InfiniBand+Clusters+with+NVIDIA+GPUs 8. DeepEP: an efficient expert-parallel communication library (and related DeepSeek-V3 Technical Report communication sections) — DeepSeek-AI (research/infra team), 2025 / 2024 https://scholar.google.com/scholar?q=DeepEP%3A+an+efficient+expert-parallel+communication+library+%28and+related+DeepSeek-V3+Technical+Report+communication+sections%29 9. Introducing OpenSHMEM: SHMEM for the PGAS Community — Barbara Chapman, Tony Curtis, Swaroop Pophale, et al., 2010 https://scholar.google.com/scholar?q=Introducing+OpenSHMEM%3A+SHMEM+for+the+PGAS+Community 10. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024 https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving 11. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 12. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models Interactive Visualization: Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing

  6. −5 d

    Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits

    This episode examines "Silicon Showdown," a study comparing Nvidia discrete-GPU and Apple unified-memory architectures for running large language models on consumer hardware, tested across model sizes from 1.5 billion to 80 billion parameters. It explains why Nvidia's VRAM Wall forces a stark trade-off between quantizing models down or offloading to slower system RAM across a PCIe bottleneck, while Apple's unified memory pool lets large models load fully without that penalty, at the cost of slower per-byte bandwidth. The discussion breaks down the competing software stacks—Nvidia's TensorRT-LLM with its new NVFP4 format and split-backend behavior, Apple's compilation-free MLX, and the cross-platform GGUF fallback from llama.cpp—and how each shapes real-world performance on metrics like time-to-first-token and tokens per joule. The episode highlights a gap in existing benchmarks like MLPerf and vLLM research, which focus on data-center throughput rather than the moment a model outgrows a single consumer GPU's memory. Listeners interested in running frontier open-weight models like Llama-3.3-70B or Qwen3-Next-80B on their own hardware will find a grounded, hardware-specific account of where each platform's approach breaks down. Sources: 1. Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference — Abdurrahman Javat, Allan Kazakov, 2026 http://arxiv.org/abs/2605.00519 2. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Frantar, Ashkboos, Hoefler, Alistarh, 2022 https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers 3. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Lin, Tang, Tang, Yang, Dang, Han, 2023 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Dettmers, Lewis, Belkada, Zettlemoyer, 2022 https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale 5. Mixtral of Experts — Jiang et al. (Mistral AI), 2024 https://scholar.google.com/scholar?q=Mixtral+of+Experts 6. Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention Interactive Visualization: Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits

  7. 4 aug.

    DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading

    This episode covers "DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch," which tackles a hidden cost in dynamic sparse KV-cache systems: the GPU-resident bookkeeping state (landmarks, reconstructed keys) used to make host-memory offloading fast can itself consume up to 64% of GPU memory — 8.5 times larger than the actual sparse KV entries it's meant to retrieve. Drawing on the lineage from H2O's heavy-hitter observation to ShadowKV's landmark-based retrieval, the discussion explains how this auxiliary overhead quietly erodes the memory savings these systems promise, with ShadowKV reaching only 6.7% of its idealized batch-size capacity on a 32-billion-parameter model. DualDecoder's proposed fix is predictive prefetching: rather than permanently parking retrieval-support state on the GPU, it predicts the next decoding step's needs one step ahead and pulls entries from host memory just in time. Listeners interested in LLM inference efficiency will find a concrete, measured account of how a fix for one memory wall can quietly build a smaller one right next to it — and a proposed way out. Sources: 1. DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch — Zuning Liang, Zhiyi Yao, Qi Chen, Yuedong Xu, Hao Dai, Zhiqiang Ding, Tongkai Yang, Jinlong Hou, Yuan Cheng, 2026 http://arxiv.org/abs/2607.26475 2. SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs — Jiaming Xu, Jiayi Pan, Hanzhen Wang, Yongkang Zhou, Jiancai Ye, Yu Wang, Guohao Dai, 2026 https://scholar.google.com/scholar?q=SpeContext%3A+Enabling+Efficient+Long-context+Reasoning+with+Speculative+Context+Sparsity+in+LLMs 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval — Di Liu, Meng Chen, Baotong Lu, et al., 2024 https://scholar.google.com/scholar?q=RetrievalAttention%3A+Accelerating+Long-Context+LLM+Inference+via+Vector+Retrieval 5. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, et al., 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 6. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+%28vLLM%29 Interactive Visualization: DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading

  8. 4 aug.

    FreeAct: Rethinking One-to-One Transforms for LLM Quantization

    This episode explores FreeAct, a new approach to quantizing large language models down to 4-bit weights and activations (W4A4), presented by researchers from the National University of Singapore, Huawei Technology, and Central South University. The discussion traces how prior methods like QuaRot and FlatQuant rely on a rigid one-to-one pairing between a rotation matrix applied to activations and its exact inverse applied to weights — an assumption that breaks down for diffusion language models, where masked and unmasked tokens have different statistical profiles, and for multimodal models mixing vision and text tokens through the same layers. The hosts unpack the outlier-channel problem that makes activation quantization so much harder than weight quantization, tracing it back to Dettmers' LLM.int8 findings, and explain how FreeAct exploits a linear-algebra insight — dubbed Proposition 1 — showing that rank-deficient activation matrices allow a whole family of transformations rather than a single exact inverse, enabling different token types to use different activation-side matrices while keeping one shared weight-side transform. It's a compelling listen for anyone tracking how quantization techniques are adapting to increasingly heterogeneous token streams in modern AI systems. Sources: 1. FreeAct: Freeing Activations for LLM Quantization — Xiaohao Liu, Xiaobo Xia, Manyi Zhang, Ji-Fu Li, Xianzhi Yu, Fei Shen, Xiu Su, See-Kiong Ng, Tat-Seng Chua, 2026 http://arxiv.org/abs/2603.01776 2. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022 https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale 3. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, Song Han, 2023 https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models 4. Atom: Low-bit Quantization for Efficient and Accurate LLM Serving — Yilong Zhao, Chien-Yu Lin, Kan Zhu, et al., 2024 https://scholar.google.com/scholar?q=Atom%3A+Low-bit+Quantization+for+Efficient+and+Accurate+LLM+Serving 5. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian Croci, et al., 2024 https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs 6. FlatQuant: Flatness Matters for LLM Quantization — Yuxuan Sun, et al., 2025 https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization 7. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, et al. (Meta AI), 2024 https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations 8. QuIP: 2-Bit Quantization of Large Language Models With Guarantees — Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa, 2023 https://scholar.google.com/scholar?q=QuIP%3A+2-Bit+Quantization+of+Large+Language+Models+With+Guarantees 9. MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization — Jiangyong Yu, Sifan Zhou, Dawei Yang, et al., 2025 https://scholar.google.com/scholar?q=MQuant%3A+Unleashing+the+Inference+Potential+of+Multimodal+Large+Language+Models+via+Static+Quantization 10. DLLMQuant: Quantizing Diffusion-based Large Language Models — Chen Xu, Dan Yang, 2025 https://scholar.google.com/scholar?q=DLLMQuant%3A+Quantizing+Diffusion-based+Large+Language+Models 11. Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models — Tianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin, Yong Guo, Yulun Zhang, 2025 https://scholar.google.com/scholar?q=Quant-dLLM%3A+Post-Training+Extreme+Low-Bit+Quantization+for+Diffusion+Large+Language+Models 12. DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs — Haokun Lin, Haobo Xu, Yichen Wu, et al., 2024 https://scholar.google.com/scholar?q=DuQuant%3A+Distributing+Outliers+via+Dual+Transformation+Makes+Stronger+Quantized+LLMs Interactive Visualization: FreeAct: Rethinking One-to-One Transforms for LLM Quantization

Om

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

Du kanske också gillar