AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 1 day ago

    Preconditioned Optimization Without the Full Matrix Cost

    This episode explores Shampoo, a 2018 optimization algorithm from Google Brain and Princeton that brings second-order curvature information to neural network training without the prohibitive cost of a full preconditioner. The discussion traces the lineage from Newton's method through AdaGrad's diagonal and full-matrix variants, explaining how Shampoo exploits the natural tensor structure of neural network weights — keeping separate small preconditioners per dimension and combining them via Kronecker products rather than materializing an impossibly large matrix. The hosts debate how much weight to put on the algorithm's convergence proofs, which rely on online convex optimization theory even though real network training is highly non-convex, concluding that the theory offers a sanity check rather than a guarantee and that empirical performance does the real work of justification. Along the way, they place Shampoo alongside K-FAC as one of the major structure-aware curvature approximations in the optimization literature. Listeners interested in the tradeoffs between cheap first-order methods and expensive second-order ones will find a clear walkthrough of how Shampoo threads that needle. Sources: 1. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 http://arxiv.org/abs/1802.09568 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Scalable Second Order Optimization for Deep Learning (Distributed Shampoo) — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2021 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning+%28Distributed+Shampoo%29 4. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 5. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 6. Adam: A Method for Stochastic Optimization — Diederik Kingma, Jimmy Ba, 2015 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 7. Decoupled Weight Decay Regularization — Ilya Loshchilov, Frank Hutter, 2017 https://scholar.google.com/scholar?q=Decoupled+Weight+Decay+Regularization 8. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 9. A Stochastic Quasi-Newton Method for Large-Scale Optimization — Richard Byrd, Samantha Hansen, Jorge Nocedal, Yoran Singer, 2016 https://scholar.google.com/scholar?q=A+Stochastic+Quasi-Newton+Method+for+Large-Scale+Optimization 10. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 11. Online Convex Programming and Generalized Infinitesimal Gradient Ascent — Martin Zinkevich, 2003 https://scholar.google.com/scholar?q=Online+Convex+Programming+and+Generalized+Infinitesimal+Gradient+Ascent 12. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012 https://scholar.google.com/scholar?q=Online+Learning+and+Online+Convex+Optimization 13. Attention Is All You Need — A. Vaswani, N. Shazeer, N. Parmar, et al., 2017 https://scholar.google.com/scholar?q=Attention+Is+All+You+Need 14. A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization — V. Gupta, T. Koren, Y. Singer, 2017 https://scholar.google.com/scholar?q=A+Unified+Approach+to+Adaptive+Regularization+in+Online+and+Stochastic+Optimization 15. Understanding Deep Learning Requires Rethinking Generalization — C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, 2017 https://scholar.google.com/scholar?q=Understanding+Deep+Learning+Requires+Rethinking+Generalization Interactive Visualization: Preconditioned Optimization Without the Full Matrix Cost

  2. 1 day ago

    Scalable Second-Order Optimization: Shampoo at Scale

    This episode explores the paper "Scalable Second Order Optimization for Deep Learning" and its introduction of Distributed Shampoo, a Kronecker-factored second-order optimizer that cuts training steps in half compared to a well-tuned Adam baseline on WMT'14 English-to-French translation, with further gains shown on BERT, Criteo click-through-rate modeling, and ResNet-50. The discussion traces the lineage from Newton's method and full-matrix AdaGrad's prohibitive O(N²)/O(N³) costs through Shampoo's Kronecker-product approximation, which replaces one massive preconditioner with smaller per-dimension matrices to make second-order optimization tractable at scale. Rival factored approaches, K-FAC and K-BFGS, are positioned as points of comparison throughout the paper. The conversation is aimed at listeners who use Adam daily but have never unpacked the mechanical distinction between first-order and second-order optimization, or the meaning of "preconditioning" itself. It's a compelling listen because second-order methods have long been dismissed as theoretically superior but practically unscalable, and this paper demonstrates real wall-clock wins across four distinct production-scale workloads rather than a single cherry-picked benchmark. Sources: 1. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 http://arxiv.org/abs/2002.09018 2. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 3. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 4. Adam: A Method for Stochastic Optimization — Diederik P. Kingma, Jimmy Ba, 2014 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 5. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization (AdaGrad) — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization+%28AdaGrad%29 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — J. Martens, R. Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 7. Distributed Second-Order Optimization using Kronecker-Factored Approximations — J. Ba, J. Martens, R. Grosse, 2017 https://scholar.google.com/scholar?q=Distributed+Second-Order+Optimization+using+Kronecker-Factored+Approximations 8. Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks — K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, S. Matsuoka, 2019 https://scholar.google.com/scholar?q=Large-Scale+Distributed+Second-Order+Optimization+Using+Kronecker-Factored+Approximate+Curvature+for+Deep+Convolutional+Neural+Networks 9. Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes — Y. You, J. Li, S. Reddi, et al. (LAMB), 2019 https://scholar.google.com/scholar?q=Large+Batch+Optimization+for+Deep+Learning%3A+Training+BERT+in+76+Minutes 10. A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes — Z. Nado, J. Gilmer, C. Shallue, R. Anil, G. Dahl, 2021 https://scholar.google.com/scholar?q=A+Large+Batch+Optimizer+Reality+Check%3A+Traditional%2C+Generic+Optimizers+Suffice+Across+Batch+Sizes 11. Limitations of the Empirical Fisher Approximation for Natural Gradient Descent — F. Kunstner, P. Hennig, L. Balles, 2019 https://scholar.google.com/scholar?q=Limitations+of+the+Empirical+Fisher+Approximation+for+Natural+Gradient+Descent Interactive Visualization: Scalable Second-Order Optimization: Shampoo at Scale

  3. 5 days ago

    Naive Test-Time Adaptation Destabilizes LLM Predictions

    This episode explores a paper proposing SCALENET, a hypernetwork-based fix for unsupervised test-time adaptation (TTA) in large language models. It examines why naive per-prompt gradient updates are unstable — a 70-billion-parameter Llama model's negative log-likelihood balloons from 2.21 to 11.49 after just five adaptation steps — and traces the problem to high-variance single-sample gradients that can't average out the way batch training does. The discussion covers the constrained "adapt-and-reset" setup used in real deployment, where models take a few unsupervised gradient steps on LoRA attention matrices per prompt before discarding the update, and explains why a single global learning rate can't work when small rates do nothing and large ones destroy the model. Listeners interested in the mechanics of on-the-fly model adaptation, LoRA-based efficient tuning, and the control-theory-like challenge of stabilizing per-layer, per-step learning rates will find the breakdown of the failure modes and the proposed hypernetwork solution especially compelling. Sources: 1. Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs — Longhuan Xu, Cunjian Chen, Feng Yin, 2026 http://arxiv.org/abs/2602.09719 2. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, Moritz Hardt, 2020 https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts 3. Tent: Fully Test-Time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2021 https://scholar.google.com/scholar?q=Tent%3A+Fully+Test-Time+Adaptation+by+Entropy+Minimization 4. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2024 https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models 5. The Surprising Effectiveness of Test-Time Training for Abstract Reasoning — Ekin Akyürek, Mehul Damani, Linlu Qiu, Han Guo, Yoon Kim, Jacob Andreas, 2024 https://scholar.google.com/scholar?q=The+Surprising+Effectiveness+of+Test-Time+Training+for+Abstract+Reasoning 6. Test-time Learning for Large Language Models — Hu, J., Zhang, Z., Chen, G., Wen, X., Shuai, C., Luo, W., Xiao, B., Li, Y., Tan, M., 2025 https://scholar.google.com/scholar?q=Test-time+Learning+for+Large+Language+Models 7. SLOT: Sample-specific Language Model Optimization at Test-time — Hu, Y., Zhang, X., Fang, X., Chen, Z., Wang, X., Zhang, H., Qi, G., 2025 https://scholar.google.com/scholar?q=SLOT%3A+Sample-specific+Language+Model+Optimization+at+Test-time 8. COME: Test-time Adaption by Conservatively Minimizing Entropy — Zhang, Q., Bian, Y., Kong, X., Zhao, P., Zhang, C., 2024 https://scholar.google.com/scholar?q=COME%3A+Test-time+Adaption+by+Conservatively+Minimizing+Entropy 9. Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models — Rannen-Triki, A., Bornschein, J., Pascanu, R., Hutter, M., et al., 2024 https://scholar.google.com/scholar?q=Revisiting+Dynamic+Evaluation%3A+Online+Adaptation+for+Large+Language+Models 10. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (MAML) — Finn, C., Abbeel, P., Levine, S., 2017 https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks+%28MAML%29 Interactive Visualization: Naive Test-Time Adaptation Destabilizes LLM Predictions

  4. 6 days ago

    Continual Learning in LLMs: Beyond Catastrophic Forgetting

    This episode explores a survey on continual learning in large language models, examining how models can be updated after pretraining without the prohibitive cost of full retraining or the risk of catastrophic forgetting — the phenomenon where new training quietly degrades performance on tasks a model previously handled well. The discussion breaks down the problem across three distinct LLM training stages (pretraining, fine-tuning, and alignment) and maps them onto three classical mitigation strategies: rehearsal-based methods that replay old data, regularization-based methods that penalize changes to critical parameters, and architecture-based methods that add task-specific capacity like adapters or LoRA modules while freezing the rest. The hosts debate the survey's core organizational claim — that structuring the literature by mechanism rather than by application domain (medical, legal, financial) offers a more useful lens for practitioners trying to borrow a specific forgetting-mitigation technique. Listeners interested in the practical tradeoffs of keeping frontier models current — especially around data that can never legally enter a pretraining corpus, like medical or financial records — will find this a grounded framing of a problem every deployed LLM eventually faces. Sources: 1. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities — Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin, 2026 http://arxiv.org/abs/2603.12658 2. Don't Stop Pretraining: Adapt Language Models to Domains and Tasks — Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith, 2020 https://scholar.google.com/scholar?q=Don%27t+Stop+Pretraining%3A+Adapt+Language+Models+to+Domains+and+Tasks 3. Simple and Scalable Strategies to Continually Pre-train Large Language Models — Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish, 2024 https://scholar.google.com/scholar?q=Simple+and+Scalable+Strategies+to+Continually+Pre-train+Large+Language+Models 4. LLaMA Pro: Progressive LLaMA with Block Expansion — Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo, 2024 https://scholar.google.com/scholar?q=LLaMA+Pro%3A+Progressive+LLaMA+with+Block+Expansion 5. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and the Code Llama team at Meta AI, 2023 https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code 6. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic 7. TIES-Merging: Resolving Interference When Merging Models — Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, Mohit Bansal, 2023 https://scholar.google.com/scholar?q=TIES-Merging%3A+Resolving+Interference+When+Merging+Models 8. Overcoming Catastrophic Forgetting in Neural Networks (EWC) — James Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks+%28EWC%29 Interactive Visualization: Continual Learning in LLMs: Beyond Catastrophic Forgetting

  5. 6 days ago

    Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

    This episode explores a paper examining what happens to continual learning problems when LLM agents shift from parametric updates to memory-augmented architectures. Rather than accepting the industry assumption that external memory sidesteps catastrophic forgetting entirely, the researchers run classic continual-learning protocols on memory-based agents and find the same core problem resurfaces in a new form — shifting from parameter capacity to context-window retrieval capacity. They identify three specific failure modes: retrieval pollution (irrelevant memories crowding the prompt), context competition (useful memories getting displaced by other retrieved items), and memory dilution (relevant material becoming harder to surface as the memory store grows). The discussion traces this argument against the history of catastrophic forgetting and prior mitigation techniques like Elastic Weight Consolidation and Gradient Episodic Memory, then explains how the paper reframes the stability-plasticity dilemma for retrieval-based systems. Listeners interested in agent design, RAG architectures, or the assumptions underlying memory-augmented LLMs will find the paper's reframing — that memory doesn't eliminate the bottleneck, it just relocates it — a useful corrective to a widely repeated industry pitch. Sources: 1. Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall https://arxiv.org/pdf/2604.27003 2. A-Mem: Agentic Memory for LLM Agents — Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang, 2025 https://scholar.google.com/scholar?q=A-Mem%3A+Agentic+Memory+for+LLM+Agents 3. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 4. How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior — Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, Zhen Xiang, 2025 https://scholar.google.com/scholar?q=How+Memory+Management+Impacts+LLM+Agents%3A+An+Empirical+Study+of+Experience-Following+Behavior 5. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su, 2025 https://scholar.google.com/scholar?q=From+RAG+to+Memory%3A+Non-Parametric+Continual+Learning+for+Large+Language+Models 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond Interactive Visualization: Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

  6. 6 days ago

    In-Place Test-Time Training Turns Fast Weights Into Online Memory

    This episode explores a new test-time training method called In-Place TTT, which repurposes the down-projection matrix inside a model's existing gated MLP as adaptable "fast weights," letting a pretrained model keep learning during inference without any architectural changes. A key innovation is replacing the reconstruction-style training target used in prior TTT approaches with an LM-aligned target built from a causal convolution over token embeddings, which the authors prove (via an induction-head theorem) actually raises the probability of the correct next token. The discussion covers how a context-parallel scan preserves causality while enabling parallel computation of these updates, and walks through benchmark results showing the method trailing a baseline at short context but pulling substantially ahead as sequence length grows, tested across Qwen3-4B, LLaMA-3.1-8B, and Qwen3-14B. The hosts also dig into an ablation showing that mid-sized chunk sizes outperform larger ones — a counterintuitive result tied to how often the fast weights get to update rather than raw parallelism — plus efficiency data showing the approach barely affects throughput or memory. It's a concrete look at how far you can push adaptive inference-time learning while reusing a model's own existing structure. Sources: 1. In-Place Test-Time Training — Guhao Feng, Shengjie Luo, Kai Hua, Ge Zhang, Di He, Wenhao Huang, Tianle Cai, 2026 http://arxiv.org/abs/2604.06169 2. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, et al., 2024 https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States 3. Test-Time Training Done Right (LaCT) — Tianyuan Zhang, Sai Bi, Yicong Hong, et al., 2025 https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right+%28LaCT%29 4. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024 https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time 5. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2020 https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories 6. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 7. LoRA: Low-Rank Adaptation of Large Language Models — Edward Hu, Yelong Shen, Phillip Wallis, et al., 2022 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models Interactive Visualization: In-Place Test-Time Training Turns Fast Weights Into Online Memory

  7. 6 days ago

    Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention

    This episode revisits Teuvo Kohonen's 1972 paper "Correlation Matrix Memories," which reframes associative memory as a hardware fault-tolerance problem rather than a representation-learning one. Kohonen builds a memory from outer-product sums of key and data vectors, then shows mathematically how much recall quality degrades when connections are randomly dropped (an "incomplete" correlation matrix memory) rather than fully wired. The discussion traces the paper's lineage against optical holography models and Steinbuch's Lernmatrix, and unpacks concepts like crosstalk and graceful degradation as information gets smeared additively across the matrix instead of stored in one fragile spot. A tangent draws — and partly disputes — a comparison between Kohonen's outer-product accumulation and the mechanics underlying modern attention, debating whether the resemblance is structural or purely coincidental given the total absence of learning or gradients in the original scheme. Listeners interested in the deep history of neural memory models and how old hardware constraints shaped ideas that echo in today's architectures will find plenty to chew on. Sources: 1. Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention https://lucidar.me/fr/neural-networks/files/1972-correlation-matrix-memories.pdf 2. Neural networks and physical systems with emergent collective computational abilities — John J. Hopfield, 1982 https://scholar.google.com/scholar?q=Neural+networks+and+physical+systems+with+emergent+collective+computational+abilities 3. Non-Holographic Associative Memory — David Willshaw, O. P. Buneman, H. Christopher Longuet-Higgins, 1969 https://scholar.google.com/scholar?q=Non-Holographic+Associative+Memory 4. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, et al., 2020 https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need 5. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021 https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers 6. A Simple Neural Network Generating an Interactive Memory — James A. Anderson, 1972 https://scholar.google.com/scholar?q=A+Simple+Neural+Network+Generating+an+Interactive+Memory 7. Representation of Associated Data by Matrix Operators — Teuvo Kohonen, Matti Ruohonen, 1973 https://scholar.google.com/scholar?q=Representation+of+Associated+Data+by+Matrix+Operators 8. Sparse Distributed Memory — Pentti Kanerva, 1988 https://scholar.google.com/scholar?q=Sparse+Distributed+Memory 9. Die Lernmatrix — K. Steinbuch, 1961 https://scholar.google.com/scholar?q=Die+Lernmatrix 10. Associative holographic memories — D. Gabor, 1969 https://scholar.google.com/scholar?q=Associative+holographic+memories 11. A class of randomly organized associative memories — T. Kohonen, 1971 https://scholar.google.com/scholar?q=A+class+of+randomly+organized+associative+memories Interactive Visualization: Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention

  8. 6 days ago

    Learning, Fast and Slow: LLMs That Adapt Without Forgetting

    This episode explores catastrophic forgetting and plasticity loss in RL-trained language models, and introduces "Fast-Slow Training," a method combining slow weight updates (RLVR) with fast in-context learning to address both. The hosts unpack the distinction between RLVR's automatic, verifiable rewards and traditional RLHF, then dig into two separate failure modes of pure RL post-training: models forgetting general competence while chasing a narrow reward signal, and a subtler loss of plasticity where updates leave models increasingly unable to absorb new tasks. Framing the two training channels as a System 1/System 2 split, the discussion centers on the paper's headline result — combining both channels reaches RL's peak accuracy with up to three times fewer samples, drifts up to seventy percent less from the base model, and preserves the capacity to learn subsequent tasks where pure RL stalls. Listeners interested in the mechanics and tradeoffs of continual learning in large language models will find a grounded walkthrough of why prompting alone hits a ceiling and why weight updates alone come with hidden costs. Sources: 1. Learning, Fast and Slow: Towards LLMs That Adapt Continually — Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal, Joseph E. Gonzalez, Matei Zaharia, Kurt Keutzer, Inderjit S Dhillon, Rishabh Agarwal, Devvrit Khatri, 2026 http://arxiv.org/abs/2605.12484 2. Loss of Plasticity in Deep Continual Learning — Shibhansh Dohare, J. Fernando Hernandez-Garcia, Qingfeng Lan, A. Rupam Mahmood, Richard S. Sutton, et al., 2024 (Nature; preprint circulated as 'Maintaining Plasticity via Continual Backprop' from 2021) https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning 3. On Warm-Starting Neural Network Training — Jordan T. Ash, Ryan P. Adams, 2020 (NeurIPS) https://scholar.google.com/scholar?q=On+Warm-Starting+Neural+Network+Training 4. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022 (ICML) https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning 5. Understanding Plasticity in Neural Networks — Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, Will Dabney, 2023 (ICML) https://scholar.google.com/scholar?q=Understanding+Plasticity+in+Neural+Networks 6. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025 https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less 7. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL) — Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S. Dhillon, David Brandfonbrener, Rishabh Agarwal, 2025 https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs+%28ScaleRL%29 8. Fine-tuning and prompt optimization: Two great steps that work better together (BetterTogether) — Dilara Soylu, Christopher Potts, Omar Khattab, 2024 https://scholar.google.com/scholar?q=Fine-tuning+and+prompt+optimization%3A+Two+great+steps+that+work+better+together+%28BetterTogether%29 9. Mitigating plasticity loss in continual reinforcement learning by reducing churn — Hongyao Tang, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Glen Berseth, 2025 https://scholar.google.com/scholar?q=Mitigating+plasticity+loss+in+continual+reinforcement+learning+by+reducing+churn 10. What can you do when you have zero rewards during RL? — Jatin Prakash, Anirudh Buvanesh, 2025 https://scholar.google.com/scholar?q=What+can+you+do+when+you+have+zero+rewards+during+RL%3F Interactive Visualization: Learning, Fast and Slow: LLMs That Adapt Without Forgetting

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like