AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 23h ago

    Continual Learning in LLMs: Beyond Catastrophic Forgetting

    This episode explores a survey on continual learning in large language models, examining how models can be updated after pretraining without the prohibitive cost of full retraining or the risk of catastrophic forgetting — the phenomenon where new training quietly degrades performance on tasks a model previously handled well. The discussion breaks down the problem across three distinct LLM training stages (pretraining, fine-tuning, and alignment) and maps them onto three classical mitigation strategies: rehearsal-based methods that replay old data, regularization-based methods that penalize changes to critical parameters, and architecture-based methods that add task-specific capacity like adapters or LoRA modules while freezing the rest. The hosts debate the survey's core organizational claim — that structuring the literature by mechanism rather than by application domain (medical, legal, financial) offers a more useful lens for practitioners trying to borrow a specific forgetting-mitigation technique. Listeners interested in the practical tradeoffs of keeping frontier models current — especially around data that can never legally enter a pretraining corpus, like medical or financial records — will find this a grounded framing of a problem every deployed LLM eventually faces. Sources: 1. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities — Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin, 2026 http://arxiv.org/abs/2603.12658 2. Don't Stop Pretraining: Adapt Language Models to Domains and Tasks — Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith, 2020 https://scholar.google.com/scholar?q=Don%27t+Stop+Pretraining%3A+Adapt+Language+Models+to+Domains+and+Tasks 3. Simple and Scalable Strategies to Continually Pre-train Large Language Models — Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish, 2024 https://scholar.google.com/scholar?q=Simple+and+Scalable+Strategies+to+Continually+Pre-train+Large+Language+Models 4. LLaMA Pro: Progressive LLaMA with Block Expansion — Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo, 2024 https://scholar.google.com/scholar?q=LLaMA+Pro%3A+Progressive+LLaMA+with+Block+Expansion 5. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and the Code Llama team at Meta AI, 2023 https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code 6. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic 7. TIES-Merging: Resolving Interference When Merging Models — Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, Mohit Bansal, 2023 https://scholar.google.com/scholar?q=TIES-Merging%3A+Resolving+Interference+When+Merging+Models 8. Overcoming Catastrophic Forgetting in Neural Networks (EWC) — James Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks+%28EWC%29 Interactive Visualization: Continual Learning in LLMs: Beyond Catastrophic Forgetting

  2. 23h ago

    Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

    This episode explores a paper examining what happens to continual learning problems when LLM agents shift from parametric updates to memory-augmented architectures. Rather than accepting the industry assumption that external memory sidesteps catastrophic forgetting entirely, the researchers run classic continual-learning protocols on memory-based agents and find the same core problem resurfaces in a new form — shifting from parameter capacity to context-window retrieval capacity. They identify three specific failure modes: retrieval pollution (irrelevant memories crowding the prompt), context competition (useful memories getting displaced by other retrieved items), and memory dilution (relevant material becoming harder to surface as the memory store grows). The discussion traces this argument against the history of catastrophic forgetting and prior mitigation techniques like Elastic Weight Consolidation and Gradient Episodic Memory, then explains how the paper reframes the stability-plasticity dilemma for retrieval-based systems. Listeners interested in agent design, RAG architectures, or the assumptions underlying memory-augmented LLMs will find the paper's reframing — that memory doesn't eliminate the bottleneck, it just relocates it — a useful corrective to a widely repeated industry pitch. Sources: 1. Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall https://arxiv.org/pdf/2604.27003 2. A-Mem: Agentic Memory for LLM Agents — Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang, 2025 https://scholar.google.com/scholar?q=A-Mem%3A+Agentic+Memory+for+LLM+Agents 3. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 4. How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior — Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, Zhen Xiang, 2025 https://scholar.google.com/scholar?q=How+Memory+Management+Impacts+LLM+Agents%3A+An+Empirical+Study+of+Experience-Following+Behavior 5. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su, 2025 https://scholar.google.com/scholar?q=From+RAG+to+Memory%3A+Non-Parametric+Continual+Learning+for+Large+Language+Models 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond Interactive Visualization: Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

  3. 23h ago

    TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization

    This episode explores TwinQuant, a 4-bit post-training quantization method for large language models that challenges a core assumption behind prior techniques like SVDQuant: that a weight matrix's important information can be captured in a small, fixed set of directions. The hosts explain how LLM weight outliers turn out to be spread across hundreds of directions rather than concentrated, forcing earlier low-rank decomposition approaches into an unwinnable tradeoff between speed and accuracy. They unpack TwinQuant's solution — learning the low-rank split itself via manifold optimization, using a true orthogonal (Stiefel manifold) rotation that folds cleanly into RMSNorm layers alongside a more flexible invertible (general linear) transform for layer-specific residual handling — plus a fused kernel designed to keep the approach fast at inference. Along the way, the conversation walks through foundational quantization vocabulary (PTQ, WxAy notation, mixed-precision splits) for listeners newer to the topic. It's a compelling listen for anyone tracking how far LLMs can be compressed without sacrificing accuracy, and why the math behind "which parts of a weight matrix matter" is more complicated than earlier compression work assumed. Sources: 1. TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization — Haodong Wang, Junjie Liu, Zicong Hong, Qianli Liu, Jian Lin, Song Guo, Xu Chen, 2026 http://arxiv.org/abs/2606.01556 2. Optimization Algorithms on Matrix Manifolds — P.-A. Absil, R. Mahony, R. Sepulchre, 2008 https://scholar.google.com/scholar?q=Optimization+Algorithms+on+Matrix+Manifolds 3. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, Igor Fedorov, et al. (Meta AI), 2024 https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations 4. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024 https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs 5. Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform — Jun Li, Fuxin Li, Sinisa Todorovic, 2020 https://scholar.google.com/scholar?q=Efficient+Riemannian+Optimization+on+the+Stiefel+Manifold+via+the+Cayley+Transform 6. SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models — Li, M., Lin, Y., Zhang, Z., Cai, T., Li, X., Guo, J., Xie, E., Meng, C., Zhu, J.-Y., Han, S., 2025 https://scholar.google.com/scholar?q=SVDQuant%3A+Absorbing+Outliers+by+Low-Rank+Components+for+4-Bit+Diffusion+Models 7. FlatQuant: Flatness Matters for LLM Quantization — Sun, Y., Liu, R., Bai, H., Bao, H., Zhao, K., Li, Y., Yu, X., Hou, L., Yuan, C., Jiang, X., et al., 2025 https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization 8. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models — Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., Luo, P., 2024 https://scholar.google.com/scholar?q=OmniQuant%3A+Omnidirectionally+Calibrated+Quantization+for+Large+Language+Models Interactive Visualization: TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization

  4. 23h ago

    Volatility Optimization Is Actually Bayesian Inference

    This episode explores Kohei Honda's tutorial and survey "Model Predictive Control via Probabilistic Inference," which unifies two decades of scattered research—path integral control, reinforcement learning theory, and variational inference—into a single coherent framework called PI-MPC. The discussion traces why classical gradient- and Hessian-based MPC solvers break down on contact-rich robotics, learned neural dynamics, or discontinuous costs, and why the resulting fallback to naive random-shooting sampling collapses under the curse of dimensionality. The core argument is that reframing sampling-based MPC as inference over a distribution of good control sequences—rather than search for a single optimum—yields dramatic gains in sample efficiency and parallelizability, with MPPI's Boltzmann-weighted, temperature-controlled posterior serving as the paper's central worked example. Along the way, the hosts debate whether "inference" is meaningfully different from optimization, tracing how entropy terms in algorithms like Soft Actor-Critic emerge naturally from the probabilistic framing rather than being added as an exploration hack. Listeners interested in robotics, control theory, or the mathematical bridges between classical control and modern probabilistic ML will find the episode's account of why this synthesis only became practical with GPU-scale parallel rollouts particularly compelling. Sources: 1. Model Predictive Control via Probabilistic Inference: A Tutorial and Survey — Kohei Honda, 2025 http://arxiv.org/abs/2511.08019v4 2. Constrained Model Predictive Control: Stability and Optimality — D. Q. Mayne, J. B. Rawlings, C. V. Rao, P. O. M. Scokaert, 2000 https://scholar.google.com/scholar?q=Constrained+Model+Predictive+Control%3A+Stability+and+Optimality 3. A Survey of Industrial Model Predictive Control Technology — S. Joe Qin, Thomas A. Badgwell, 2003 https://scholar.google.com/scholar?q=A+Survey+of+Industrial+Model+Predictive+Control+Technology 4. Model Predictive Control: Theory and Practice — A Survey — Carlos E. Garcia, David M. Prett, Manfred Morari, 1989 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey 5. Model Predictive Path Integral Control using Covariance Variable Importance Sampling — Grady Williams, Andrew Aldrich, Evangelos A. Theodorou, 2015 https://scholar.google.com/scholar?q=Model+Predictive+Path+Integral+Control+using+Covariance+Variable+Importance+Sampling 6. Information Theoretic MPC for Model-Based Reinforcement Learning — Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James M. Rehg, Byron Boots, Evangelos A. Theodorou, 2017 https://scholar.google.com/scholar?q=Information+Theoretic+MPC+for+Model-Based+Reinforcement+Learning 7. Robust Sampling Based Model Predictive Control with Sparse Objective Information — Grady Williams, Brian Goldfain, Paul Drews, Kamil Saigol, James M. Rehg, Evangelos A. Theodorou, 2018 https://scholar.google.com/scholar?q=Robust+Sampling+Based+Model+Predictive+Control+with+Sparse+Objective+Information 8. Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review — Sergey Levine, 2018 https://scholar.google.com/scholar?q=Reinforcement+Learning+and+Control+as+Probabilistic+Inference%3A+Tutorial+and+Review 9. Robot Trajectory Optimization using Approximate Inference — Marc Toussaint, 2009 https://scholar.google.com/scholar?q=Robot+Trajectory+Optimization+using+Approximate+Inference 10. Optimal Control as a Graphical Model Inference Problem — Hilbert J. Kappen, Vicenç Gómez, Manfred Opper, 2012 https://scholar.google.com/scholar?q=Optimal+Control+as+a+Graphical+Model+Inference+Problem 11. Variational Inference: A Review for Statisticians — David M. Blei, Alp Kucukelbir, Jon D. McAuliffe, 2017 https://scholar.google.com/scholar?q=Variational+Inference%3A+A+Review+for+Statisticians 12. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013 https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes 13. An Introduction to Variational Methods for Graphical Models — Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, Lawrence K. Saul, 1999 https://scholar.google.com/scholar?q=An+Introduction+to+Variational+Methods+for+Graphical+Models 14. Predictive Sampling: Real-Time Behaviour Synthesis with MuJoCo — Taylor Howell, Nimrod Gileadi, Saran Tunyasuvunakool, Kevin Zakka, Tom Erez, Yuval Tassa, 2022 https://scholar.google.com/scholar?q=Predictive+Sampling%3A+Real-Time+Behaviour+Synthesis+with+MuJoCo 15. STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation — Mohak Bhardwaj, Balakumar Sundaralingam, Arsalan Mousavian, Nathan D. Ratliff, Dieter Fox, Fabio Ramos, Byron Boots, 2021 https://scholar.google.com/scholar?q=STORM%3A+An+Integrated+Framework+for+Fast+Joint-Space+Model-Predictive+Control+for+Reactive+Manipulation 16. Information-Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving — Grady Williams, Paul Drews, Brian Goldfain, James M. Rehg, Evangelos A. Theodorou, 2018 https://scholar.google.com/scholar?q=Information-Theoretic+Model+Predictive+Control%3A+Theory+and+Applications+to+Autonomous+Driving 17. Model-Based Diffusion for Trajectory Optimization — Chaoyi Pan, Zeji Yi, Guanya Shi, Guannan Qu, 2024 https://scholar.google.com/scholar?q=Model-Based+Diffusion+for+Trajectory+Optimization 18. TD-MPC2: Scalable, Robust World Models for Continuous Control — Nicklas Hansen, Hao Su, Xiaolong Wang, 2023 https://scholar.google.com/scholar?q=TD-MPC2%3A+Scalable%2C+Robust+World+Models+for+Continuous+Control 19. Recent Advances in Path Integral Control for Trajectory Optimization: An Overview in Theoretical and Algorithmic Perspectives — Muhammad Kazim, Jungee Hong, Min-Gyeom Kim, Kwang-Ki K. Kim, 2024 https://scholar.google.com/scholar?q=Recent+Advances+in+Path+Integral+Control+for+Trajectory+Optimization%3A+An+Overview+in+Theoretical+and+Algorithmic+Perspectives Interactive Visualization: Volatility Optimization Is Actually Bayesian Inference

  5. 3d ago

    TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

    This episode explores TFGN, an architectural approach to continual pre-training of large language models that claims to solve catastrophic forgetting without four common crutches: replay buffers, task identifiers, small-scale toy benchmarks, and external penalty terms like Fisher-information regularization. The hosts trace the lineage of the forgetting problem back to 1989, explain why popular fixes like LoRA-based parameter-efficient fine-tuning don't actually address forgetting (they just shrink the blast radius), and why classic regularization methods like Elastic Weight Consolidation break down at billion-parameter scale. They also clarify why long-context windows and prompt-based knowledge aren't a substitute for genuinely updating model weights on massive, unbounded corpora like full codebases or legal archives. The conversation lays out TFGN's core mechanism as a dense, input-conditioned overlay operating inside each transformer block, contrasting it with sparse mixture-of-experts routing, and sets up backward transfer as the key metric for measuring whether old knowledge survives new training. Listeners interested in how production LLMs might eventually absorb new domains without expensive retraining or fragile adapter stacking will find the framing of this open problem sharply drawn. Sources: 1. TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale — Anurup Ganguli, 2026 http://arxiv.org/abs/2605.15053 2. Overcoming catastrophic forgetting in neural networks (EWC) — J. Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+catastrophic+forgetting+in+neural+networks+%28EWC%29 3. Loss of plasticity in deep continual learning — S. Dohare et al., 2024, Nature https://scholar.google.com/scholar?q=Loss+of+plasticity+in+deep+continual+learning 4. Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning — O. Y. L. Imanov, 2026, arXiv:2601.18699 https://scholar.google.com/scholar?q=Mechanistic+Analysis+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-Tuning 5. Examining Forgetting in Continual Pre-training of Aligned Large Language Models — C.-A. Li and H.-Y. Lee, 2024, arXiv:2401.03129 https://scholar.google.com/scholar?q=Examining+Forgetting+in+Continual+Pre-training+of+Aligned+Large+Language+Models 6. Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models — I. Abbes, G. Subbaraj, M. Riemer, et al., 2025, arXiv:2508.01908 https://scholar.google.com/scholar?q=Revisiting+Replay+and+Gradient+Alignment+for+Continual+Pre-Training+of+Large+Language+Models 7. Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought — Y. Zhang, B. Tang, T. Ju, S. Duan, G. Liu, 2025, arXiv:2512.21711 https://scholar.google.com/scholar?q=Do+Latent+Tokens+Think%3F+A+Causal+and+Adversarial+Analysis+of+Chain-of-Continuous-Thought Interactive Visualization: TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

  6. 4d ago

    SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

    This episode explores SnapStream, a technique from SambaNova Systems for compressing KV caches during long-sequence LLM decoding on dataflow accelerators, demonstrated at production scale with a 671-billion-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens per second. The discussion covers why established training-free KV cache eviction methods like SnapKV and StreamingLLM have struggled to reach real deployments despite promising accuracy results: continuous batching makes it unclear when to trigger compression across requests at different lifecycle stages, and static-graph compilers used by dataflow accelerators can't easily accommodate the dynamic, variable-shaped operations that standard compression implementations rely on. It explains how SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window approach into a single fixed-size cache, splitting sequences into sink tokens, recent tokens, and a compressed middle section during prefill. The conversation is grounded in fundamentals—clarifying the prefill/decode split, why decode is memory-bound, and what makes dataflow accelerators architecturally different from GPUs—making it accessible to listeners unfamiliar with KV cache mechanics while still delivering a specific, hardware-grounded engineering story rather than a purely algorithmic one. Sources: 1. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, Evgenii Iuliugin, Magnus Vesterlund, Christian Häggström, Guangtao Wang, Shubhangi Upasani, Ayush Sachdeva, Rui Li, Faline Fu, Chen Wu, Ayesha Siddiqua, John Long, Tuowen Zhao, Matheen Musaddiq, Håkan Zeffer, Yun Du, Mingran Wang, Qinghua Li, Bo Li, Urmish Thakker, Raghu Prabhakar, 2025 http://arxiv.org/abs/2511.03092 2. Plasticine: A Reconfigurable Architecture For Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+For+Parallel+Patterns 3. SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — Raghu Prabhakar and SambaNova Systems architecture team, 2024 https://scholar.google.com/scholar?q=SN40L%3A+Scaling+the+AI+Memory+Wall+with+Dataflow+and+Composition+of+Experts 4. Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, et al. (Groq), 2020 https://scholar.google.com/scholar?q=Think+Fast%3A+A+Tensor+Streaming+Processor+%28TSP%29+for+Accelerating+Deep+Learning+Workloads 5. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 6. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, S. Han, 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, M. Sun, 2024 https://scholar.google.com/scholar?q=InfLLM%3A+Training-Free+Long-Context+Extrapolation+with+an+Efficient+Context+Memory 9. DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025 https://scholar.google.com/scholar?q=DeepSeek-V3.2-Exp%3A+Boosting+Long-Context+Efficiency+with+DeepSeek+Sparse+Attention 10. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha, D. Zinsley, E. Liu, W. Tennien, A. Rudra, J. Zou, A. Mirhoseini, C. Re, 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study 11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — G. Wang, S. Upasani, C. Wu, D. Gandhi, J. Li, C. Hu, B. Li, U. Thakker, 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+%28SAGE-KV%29 Interactive Visualization: SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

  7. 4d ago

    StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

    This episode explores StrataCL, a fabric-native communication library from researchers at Peking University, ICT-CAS, UCAS, Shanghai Jiao Tong University, and Huawei, tested on Huawei's CloudMatrix384 supernode. The discussion centers on how communication overhead — which the paper puts at 30-45% of end-to-end time in distributed LLM training and up to 50% at scale — can be cut by giving collectives and MoE routing true zero-copy access to application buffers on unified-memory fabrics, without breaking compatibility with frameworks like PyTorch and SGLang. A key insight is why buffer-centric libraries like NCCL and HCCL fall short even on fast unified-address fabrics, and how MoE dispatch/combine traffic exposes the limits of naive redesigns. The core technical contribution is registration-on-allocation: exploiting the multi-second gap between physical memory allocation and first use by a communication operator to move registration off the critical path entirely, asynchronously, the moment memory is mapped. The result is a 1.4x iteration-time speedup on a 512-die production training run with no changes to the model, optimizer, or data — pure systems engineering payoff. Sources: 1. StrataCL: Fabric-Native Communication Library for Production Supernodes — Tiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang, 2026 http://arxiv.org/abs/2607.26444 2. U-Net: A User-Level Network Interface for Parallel and Distributed Computing — Thorsten von Eicken, Anindya Basu, Vineet Buch, Werner Vogels, 1995 https://scholar.google.com/scholar?q=U-Net%3A+A+User-Level+Network+Interface+for+Parallel+and+Distributed+Computing 3. Design Guidelines for High Performance RDMA Systems — Anuj Kalia, Michael Kaminsky, David G. Andersen, 2016 https://scholar.google.com/scholar?q=Design+Guidelines+for+High+Performance+RDMA+Systems 4. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, Ion Stoica, 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML 5. NVSHMEM (GPU-initiated, PGAS-style one-sided communication library) — NVIDIA (library/runtime, not a single academic paper), 2016 (initial release, iterated since) https://scholar.google.com/scholar?q=NVSHMEM+%28GPU-initiated%2C+PGAS-style+one-sided+communication+library%29 6. Collective Communication for 100k+ GPUs — Min Si, Pavan Balaji, Yongzhou Chen, et al., 2025 https://scholar.google.com/scholar?q=Collective+Communication+for+100k%2B+GPUs 7. SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading — Xingyi Li, Yadong Liu, Xiaojie Huang, et al., 2026 (NSDI 26) https://scholar.google.com/scholar?q=SwiftEP%3A+Accelerating+MoE+Inference+with+Buffer+Fusion+and+TMA+Offloading 8. PyTorch Symmetric Memory / NVSHMEM-style same-VA mirrored buffers — PyTorch Team / NVIDIA (NVSHMEM), 2024-2025 https://scholar.google.com/scholar?q=PyTorch+Symmetric+Memory+%2F+NVSHMEM-style+same-VA+mirrored+buffers 9. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 (SOSP 23) https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention Interactive Visualization: StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

  8. 5d ago

    MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

    This episode explores a challenge to conventional wisdom in parameter-efficient fine-tuning, examining a method called MiCA that inverts the logic behind LoRA (Low-Rank Adaptation). Rather than letting trainable weight-update matrices drift freely, as standard LoRA does, MiCA deliberately anchors one matrix to the minor singular-value directions of a weight matrix — the low-energy, rarely-used "corners" that classical compression theory says to discard — leaving those directions free for new knowledge rather than overwriting the dominant, pretrained-heavy subspace. The discussion traces the technique's lineage through SVD, the Eckart-Young-Mirsky theorem, PiSSA's SVD-based initialization, and Minor Component Analysis, framing MiCA's core bet: catastrophic forgetting during fine-tuning may stem from cramming new information into already-saturated high-energy directions. Listeners interested in the mechanics of efficient model adaptation, knowledge editing, and where the field's assumptions about "useless" weight-matrix structure might be wrong will find the debate over whether this is a genuine architectural insight or a narrower refinement of existing PEFT ideas especially engaging. Sources: 1. MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA https://arxiv.org/pdf/2604.01694 2. The Approximation of One Matrix by Another of Lower Rank — Carl Eckart, Gale Young, 1936 https://scholar.google.com/scholar?q=The+Approximation+of+One+Matrix+by+Another+of+Lower+Rank 3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models 4. PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models — Fanxu Meng, Zhaohui Wang, Muhan Zhang, 2024 https://scholar.google.com/scholar?q=PiSSA%3A+Principal+Singular+Values+and+Singular+Vectors+Adaptation+of+Large+Language+Models 5. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning — Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, Tuo Zhao, 2023 https://scholar.google.com/scholar?q=AdaLoRA%3A+Adaptive+Budget+Allocation+for+Parameter-Efficient+Fine-Tuning 6. SOMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning — Seokju Yun, Seunghye Chae, Dongheon Lee, Youngmin Ro, 2025 https://scholar.google.com/scholar?q=SOMA%3A+Singular+Value+Decomposed+Minor+Components+Adaptation+for+Domain+Generalizable+Representation+Learning 7. Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-Tuning — Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh, 2026 https://scholar.google.com/scholar?q=Learning+Rate+Matters%3A+Vanilla+LoRA+May+Suffice+for+LLM+Fine-Tuning 8. DoRA: Weight-Decomposed Low-Rank Adaptation — Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen, 2024 https://scholar.google.com/scholar?q=DoRA%3A+Weight-Decomposed+Low-Rank+Adaptation 9. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 10. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic Interactive Visualization: MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

Ratings & Reviews

3.7
out of 5
3 Ratings

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like