AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 1d ago

    AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill

    This episode explores STEEL, a sparsity-aware fused attention design for running long-sequence prefill inference efficiently on AMD's XDNA neural processing unit. The discussion contrasts spatial-dataflow NPU architectures, where compute tiles are explicitly scheduled with no dynamic cache management, against GPU SIMT execution, and explains how the causal attention mask creates load imbalance that a fixed pipeline can't easily absorb the way a GPU scheduler can. Building on FlashAttention-2's tiling and online-softmax approach, the paper restructures the computation into a three-stage pipeline across dedicated compute cores to address that imbalance directly. The hosts walk through why this matters for on-device AI agents that need low latency, privacy, and battery efficiency without offloading to cloud GPUs. Reported results include over 9.5x latency reduction versus prior state-of-the-art NPU implementations, over 9x energy savings against a CPU baseline, and more than 22x speedup over a naive layer-by-layer approach. Sources: 1. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU — Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini, 2026 http://arxiv.org/abs/2607.09385v1 2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022 https://scholar.google.com/scholar?q=FlashAttention:+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 3. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks — Yu-Hsin Chen, Joel Emer, Vivienne Sze, 2016 https://scholar.google.com/scholar?q=Eyeriss:+An+Energy-Efficient+Reconfigurable+Accelerator+for+Deep+Convolutional+Neural+Networks 4. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 5. Plasticine: A Reconfigurable Architecture for Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine:+A+Reconfigurable+Architecture+for+Parallel+Patterns 6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 7. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Ainslie et al., 2023 https://scholar.google.com/scholar?q=GQA:+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 8. Efficiently Scaling Transformer Inference — Pope et al., 2022 https://scholar.google.com/scholar?q=Efficiently+Scaling+Transformer+Inference 9. FlashDecoding++: Faster Large Language Model Inference on GPUs — Hong et al., 2024 https://scholar.google.com/scholar?q=FlashDecoding++:+Faster+Large+Language+Model+Inference+on+GPUs 10. NITRO: LLM Inference on Intel Laptop NPUs — Fei and Abdelfattah, 2024 https://scholar.google.com/scholar?q=NITRO:+LLM+Inference+on+Intel+Laptop+NPUs Interactive Visualization: AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill

  2. 1d ago

    Cost-Aware Speculative Decoding for Mixture-of-Experts Models

    This episode explores how speculative decoding — the standard trick for speeding up autoregressive LLM inference by having a cheap draft model propose tokens for batch verification — breaks down when applied to Mixture-of-Experts models. The discussion traces why decoding is memory-bandwidth bound rather than compute bound, then shows how MoE routing decouples a draft token's acceptance probability from its actual verification cost: tokens that route to disjoint experts (termed "expert scattering") force costly extra weight fetches even when a confidence-only selector rates them highly. The paper introduces EcoSpec, a cost-aware draft selector that accounts for expert-loading overhead rather than optimizing acceptance length (alpha) alone, and the hosts examine tradeoffs in Table 1 where EcoSpec sacrifices a small amount of acceptance probability on models like Qwen3 and GPT-OSS in exchange for reduced memory traffic. Listeners interested in LLM inference serving, hardware-aware systems design, or the practical limits of applying dense-model optimizations to sparse architectures will find the episode's reframing of a three-year-old assumption particularly compelling. Sources: 1. Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts — Jincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu, 2026 http://arxiv.org/abs/2607.12696 2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 3. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024 https://scholar.google.com/scholar?q=EAGLE-2:+Faster+Inference+of+Language+Models+with+Dynamic+Draft+Trees 4. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Tri Dao, et al., 2024 https://scholar.google.com/scholar?q=Medusa:+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads 5. Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models — Keisuke Kamahori, Yile Gu, Kan Zhu, Baris Kasikci, 2024 https://scholar.google.com/scholar?q=Fiddler:+CPU-GPU+Orchestration+for+Fast+Inference+of+Mixture-of-Experts+Models 6. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Li, Y., Wei, F., Zhang, C., Zhang, H., 2026 https://scholar.google.com/scholar?q=EAGLE-3:+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test 7. MoE-Spec: Expert Budgeting for Efficient Speculative Decoding — McDanel, B., Li, S., Surineni, S., Khaitan, H., 2026 https://scholar.google.com/scholar?q=MoE-Spec:+Expert+Budgeting+for+Efficient+Speculative+Decoding 8. SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference — Chen, L., Wen, Z., Wu, T., Zhang, X., Wu, C., 2025 https://scholar.google.com/scholar?q=SP-MoE:+Speculative+Decoding+and+Prefetching+for+Accelerating+MoE-based+Model+Inference 9. MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts — Wang, W., Liu, J., Hou, X., Xia, X., Tang, P., Zhang, M., Li, C., Guo, M., 2025 https://scholar.google.com/scholar?q=MoE-SpeQ:+Speculative+Quantized+Decoding+with+Proactive+Expert+Prefetching+and+Offloading+for+Mixture-of-Experts 10. Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding (GTO) — Hu, S., Li, J., Lu, Z., Zhou, P., 2026 https://scholar.google.com/scholar?q=Bridging+Draft+Policy+Misalignment:+Group+Tree+Optimization+for+Speculative+Decoding+(GTO) 11. MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache — Xue, L., Fu, Y., Lu, Z., Mai, L., Marina, M., 2025 https://scholar.google.com/scholar?q=MoE-Infinity:+Efficient+MoE+Inference+on+Personal+Machines+with+Sparsity-Aware+Expert+Cache 12. A Survey on Inference Optimization Techniques for Mixture of Experts Models — Liu, J., Tang, P., Wang, W., Ren, Y., Hou, X., Heng, P.A., Guo, M., Li, C., 2026 https://scholar.google.com/scholar?q=A+Survey+on+Inference+Optimization+Techniques+for+Mixture+of+Experts+Models Interactive Visualization: Cost-Aware Speculative Decoding for Mixture-of-Experts Models

  3. 1d ago

    HyperOffload's Scheduling Claims Under Scrutiny

    This episode examines HyperOffload, a graph-driven scheduling system from Shanghai Jiao Tong University and Huawei that shifts LLM memory offload and prefetch decisions from a reactive runtime into compile-time graph scheduling for terabyte-scale "SuperNode" hardware. The hosts scrutinize the paper's headline 26% peak memory reduction, arguing it's largely definitional since it comes from offloading the entire KV cache in one configuration, while pointing to the defragmentation results (57 stalls eliminated) and bandwidth-robustness curves as the figures that actually demonstrate the scheduler's value. They flag a notable gap: the paper's motivating anecdote about a 2.7x slowdown from reactive prefetching is never directly retested against HyperOffload, leaving its central justification unconfirmed. The discussion also surfaces missing citations to ZeRO-Infinity and a lack of engagement with PagedAttention as a competing paradigm, plus the fact that all results are confined to Ascend NPUs and MindSpore with no evidence of portability to CUDA or PyTorch. Listeners interested in memory management for large-scale LLM serving and training will find a sharp critique of how benchmark framing can overstate a system's true contribution. Sources: 1. HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures — Fangxin Liu, Qinghua Zhang, Hanjing Shen, Zhibo Liang, Li Jiang, Haibing Guan, Chong Bao, Xuefeng Jin, 2026 http://arxiv.org/abs/2602.00748 2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity:+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 3. Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization — Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Joseph Gonzalez, Kurt Keutzer, Ion Stoica, 2020 (MLSys) https://scholar.google.com/scholar?q=Checkmate:+Breaking+the+Memory+Wall+with+Optimal+Tensor+Rematerialization 4. AutoTM: Automatic Tensor Movement in Heterogeneous Memory Systems using Integer Linear Programming — Michael Hildebrand, Jawad Khan, Sanjeev Trika, Jason Lowe-Power, Venkatesh Akella, 2020 (ASPLOS) https://scholar.google.com/scholar?q=AutoTM:+Automatic+Tensor+Movement+in+Heterogeneous+Memory+Systems+using+Integer+Linear+Programming 5. G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations — Haoyang Zhang, Yirui Zhou, Yuqi Xue, Yiqi Liu, Jian Huang, 2023 (MICRO) https://scholar.google.com/scholar?q=G10:+Enabling+An+Efficient+Unified+GPU+Memory+and+Storage+Architecture+with+Smart+Tensor+Migrations 6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 7. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024 https://scholar.google.com/scholar?q=Mooncake:+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 8. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan et al. (DeepSeek-AI), 2025 https://scholar.google.com/scholar?q=Native+Sparse+Attention:+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention Interactive Visualization: HyperOffload's Scheduling Claims Under Scrutiny

  4. 1d ago

    Parcae: Stabilizing Looped Language Models with Control Theory

    This episode explores Parcae, a paper on scaling laws for stable looped language models, where instead of stacking distinct transformer layers, a single block is applied repeatedly through the residual stream — echoing Universal Transformers and ALBERT's weight-tying but tackling the training instability that has historically plagued the approach. The discussion centers on reframing looped inference as a linear time-invariant dynamical system, showing that the spectral norm of the transition matrix A determines whether the residual stream stays bounded or explodes exponentially — turning a mysterious loss-spike failure mode into a measurable, checkable quantity. It also covers how the authors diagnose this concretely by examining the eigenvalues of A (contrasting how different prelude-embedding injection methods affect stability), and how they extend Chinchilla-style isoFLOP curve-fitting with a third axis — recurrence depth — to find the FLOP-optimal number of loops at a given compute budget. Listeners interested in efficient inference, edge deployment, or the mechanics of why prior recurrent-depth models like RDM needed fragile tuning will find the control-theory framing a clarifying, math-grounded alternative to typical trial-and-error architecture papers. Sources: 1. Parcae: Scaling Laws For Stable Looped Language Models — Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick, Daniel Y. Fu, 2026 http://arxiv.org/abs/2604.12946 2. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018 https://scholar.google.com/scholar?q=Universal+Transformers 3. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations — Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, 2019 https://scholar.google.com/scholar?q=ALBERT:+A+Lite+BERT+for+Self-supervised+Learning+of+Language+Representations 4. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025 https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning:+A+Recurrent+Depth+Approach 5. Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA — Sangmin Bae et al., 2024 https://scholar.google.com/scholar?q=Relaxed+Recursive+Transformers:+Effective+Parameter+Sharing+with+Layer-wise+LoRA 6. A Proposal on Machine Learning via Dynamical Systems — Weinan E, 2017 https://scholar.google.com/scholar?q=A+Proposal+on+Machine+Learning+via+Dynamical+Systems 7. Neural Ordinary Differential Equations — Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, David Duvenaud, 2018 https://scholar.google.com/scholar?q=Neural+Ordinary+Differential+Equations 8. Stable Architectures for Deep Neural Networks — Eldad Haber, Lars Ruthotto, 2017 https://scholar.google.com/scholar?q=Stable+Architectures+for+Deep+Neural+Networks 9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 https://scholar.google.com/scholar?q=Mamba:+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces 10. Transformers are SSMs: Generalized Models and Efficient Algorithms through Structured State Space Duality — Tri Dao, Albert Gu, 2024 https://scholar.google.com/scholar?q=Transformers+are+SSMs:+Generalized+Models+and+Efficient+Algorithms+through+Structured+State+Space+Duality 11. Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation — Sangmin Bae, Yujin Kim, Reza Bayat, et al., 2025 https://scholar.google.com/scholar?q=Mixture-of-Recursions:+Learning+Dynamic+Recursive+Depths+for+Adaptive+Token-Level+Computation 12. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu, Zixuan Wang, Kai Hua, et al., 2025 https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models 13. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence — Sean McLeish, Ang Li, John Kirchenbauer, et al., 2025 https://scholar.google.com/scholar?q=Teaching+Pretrained+Language+Models+to+Think+Deeper+with+Retrofitted+Recurrence 14. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi, 2025 https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts:+On+the+Power+of+Looped+Transformers Interactive Visualization: Parcae: Stabilizing Looped Language Models with Control Theory

  5. 1d ago

    SoundnessBench: Exposing AI Reviewers' Blind Spots

    This episode examines SoundnessBench, a new benchmark testing whether frontier LLMs can judge the underlying soundness of a research proposal before any experiments are run, rather than just executing and scoring completed work like prior agent benchmarks (MLE-Bench, PaperBench, InnovatorBench). Built from 1,099 ICLR proposals labeled with reviewers' soundness sub-scores rather than acceptance outcomes, the benchmark found that twelve frontier models produced a 74% false-positive rate — repeatedly rating flawed proposals as sound. The hosts debate whether this stems from a sycophancy-style bias inherited from RLHF training, pointing to a striking result where switching to "aggressive" fault-hunting prompts flips the same models' verdicts on the same proposals, suggesting the failure is about framing sensitivity rather than missing domain knowledge. The discussion lands on why this matters for autonomous AI research agents: an unreliable judge sitting at the "first gate" risks industrializing well-executed experiments built on dead-on-arrival ideas. Sources: 1. SoundnessBench: Exposing AI Reviewers' Blind Spots https://arxiv.org/pdf/2605.30329 2. Discovering Language Model Behaviors with Model-Written Evaluations — Ethan Perez, Sam Ringer, Kamile Lukosiute, et al. (Anthropic), 2022 https://scholar.google.com/scholar?q=Discovering+Language+Model+Behaviors+with+Model-Written+Evaluations 3. Towards Understanding Sycophancy in Language Models — Mrinank Sharma, Meg Tong, Tomasz Korbak, et al. (Anthropic, with academic collaborators), 2023 https://scholar.google.com/scholar?q=Towards+Understanding+Sycophancy+in+Language+Models 4. Simple Synthetic Data Reduces Sycophancy in Large Language Models — Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, Quoc V. Le (Google DeepMind / Google Brain), 2023 https://scholar.google.com/scholar?q=Simple+Synthetic+Data+Reduces+Sycophancy+in+Large+Language+Models 5. Prompt Sensitivity Evaluations of Large Language Models — Kate Elkins, Jon Chun (and related follow-on prompt-robustness studies, e.g. Geng et al.), 2025 https://scholar.google.com/scholar?q=Prompt+Sensitivity+Evaluations+of+Large+Language+Models 6. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers — Chenglei Si, Diyi Yang, Tatsunori Hashimoto, 2025 https://scholar.google.com/scholar?q=Can+LLMs+Generate+Novel+Research+Ideas?+A+Large-Scale+Human+Study+with+100++NLP+Researchers 7. The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas — Chenglei Si, Tatsunori Hashimoto, Diyi Yang, 2025 https://scholar.google.com/scholar?q=The+Ideation-Execution+Gap:+Execution+Outcomes+of+LLM-Generated+versus+Human+Research+Ideas 8. The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search — Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, David Ha, 2025 https://scholar.google.com/scholar?q=The+AI+Scientist-v2:+Workshop-Level+Automated+Scientific+Discovery+via+Agentic+Tree+Search 9. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, et al. (OpenAI), 2025 https://scholar.google.com/scholar?q=PaperBench:+Evaluating+AI's+Ability+to+Replicate+AI+Research Interactive Visualization: SoundnessBench: Exposing AI Reviewers' Blind Spots

  6. 1d ago

    Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks

    This episode explores how inference-time scaling breaks down once large language models shift from short chat responses to long chain-of-thought reasoning, drawing on Micron and Argonne National Laboratory's research spanning models from 8 billion to 671 billion parameters. It explains the divide between compute-bound prefill and bandwidth-bound decode phases, and how reasoning traces exceeding ten thousand tokens push systems into a "capacity-bound" regime where the KV cache — not raw FLOPs — becomes the limiting resource. The discussion contrasts three parallelism strategies (data, tensor, and pipeline) and shows why data parallelism, the industry default, hits a capacity wall under reasoning workloads even though it remains optimal for short prompts. It also covers how architectural choices like Grouped-Query Attention versus DeepSeek-R1's Mixture-of-Experts design and Multi-Head Latent Attention change how much cache pressure a model generates per token. Listeners interested in the practical engineering tradeoffs behind serving reasoning models at scale will find concrete guidance on when each parallelism strategy actually wins, backed by measurements on an 8x H200 NVLink node. Sources: 1. Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles — Moiz Arif, Avinash Maurya, Sudharshan Vazhkudai, Bogdan Nicolae, 2026 http://arxiv.org/abs/2605.19775 2. PyTorch Distributed: Experiences on Accelerating Data Parallel Training — Shen Li, Yanli Zhao, Rohan Varma, et al. (Meta AI / PyTorch team), 2020 https://scholar.google.com/scholar?q=PyTorch+Distributed:+Experiences+on+Accelerating+Data+Parallel+Training 3. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models — Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He (Microsoft), 2020 https://scholar.google.com/scholar?q=ZeRO:+Memory+Optimizations+Toward+Training+Trillion+Parameter+Models 4. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al. (UC Berkeley), 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+(vLLM) 5. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun (Seoul National University / FriendliAI), 2022 https://scholar.google.com/scholar?q=Orca:+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models 6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe:+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 7. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini, 2024 https://scholar.google.com/scholar?q=Splitwise:+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 8. Llumnix: Dynamic Scheduling for Large Language Model Serving — Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, Wei Lin, 2024 https://scholar.google.com/scholar?q=Llumnix:+Dynamic+Scheduling+for+Large+Language+Model+Serving 9. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar, 2024 https://scholar.google.com/scholar?q=vAttention:+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention 10. Efficient Memory Management for Large Language Model Serving with PagedAttention (already cited [22]) — cross-check against KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+(already+cited+[22])+—+cross-check+against+KIVI:+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache Interactive Visualization: Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks

  7. 2d ago

    Deep Native Structural Reasoning for Proteins, Molecules, and Crystals

    This episode explores SciReasoner, a 29-author foundation model from Shanghai AI Laboratory and collaborators, designed to reason natively over protein, molecule, and crystal structures rather than flattened text descriptions. The discussion breaks down why standard sub-word tokenizers (like BPE) mangle chemical structures — shattering a molecule's SMILES string into 31 largely meaningless fragments — and how SciReasoner instead uses domain-specific tokenizers (Foldseek's 3Di for protein geometry, SLICES for crystals, ConfSeq for molecular conformations) to compress the same molecule into 14 tokens that preserve real structural meaning. The hosts examine retrosynthesis as a key test domain, tracing its roots to E.J. Corey's Nobel-winning "disconnection" framework, and frame the model's core claim: producing traceable reasoning grounded in addressable structural evidence instead of an opaque black-box score. Listeners interested in whether a single unified model can genuinely bridge protein biology, chemistry, and materials science — and whether its transparency claims hold up under scrutiny — will find the episode's skeptical, formalism-first approach compelling. Sources: 1. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning — Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai, 2026 http://arxiv.org/abs/2607.07708 2. Planning Chemical Syntheses with Deep Neural Networks and Symbolic AI — Marwin H. S. Segler, Mike Preuss, Mark P. Waller, 2018 https://scholar.google.com/scholar?q=Planning+Chemical+Syntheses+with+Deep+Neural+Networks+and+Symbolic+AI 3. Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction — Philippe Schwaller, Teodoro Laino, et al., 2019 https://scholar.google.com/scholar?q=Molecular+Transformer:+A+Model+for+Uncertainty-Calibrated+Chemical+Reaction+Prediction 4. A Graph to Graphs Framework for Retrosynthesis Prediction — Chence Shi, Minkai Xu, Hongyu Guo, Ming Zhang, Jian Tang, 2020 https://scholar.google.com/scholar?q=A+Graph+to+Graphs+Framework+for+Retrosynthesis+Prediction 5. Computer-Assisted Retrosynthesis Based on Molecular Similarity — Connor W. Coley, Luke Rogers, William H. Green, Klavs F. Jensen, 2017 https://scholar.google.com/scholar?q=Computer-Assisted+Retrosynthesis+Based+on+Molecular+Similarity 6. RSGPT (unnamed full title, cited as [31]) — Not given in excerpt, cited as prior template-free SOTA https://scholar.google.com/scholar?q=RSGPT+(unnamed+full+title,+cited+as+[31]) 7. Fast and accurate protein structure search with Foldseek — van Kempen et al., 2023/2024 https://scholar.google.com/scholar?q=Fast+and+accurate+protein+structure+search+with+Foldseek 8. SLICES: a simplified line-input crystal-encoding system — Xiao et al., cited as [85] https://scholar.google.com/scholar?q=SLICES:+a+simplified+line-input+crystal-encoding+system 9. ConfSeq: conformation-aware molecular sequence representation — Xiong et al., cited as [58] https://scholar.google.com/scholar?q=ConfSeq:+conformation-aware+molecular+sequence+representation 10. DAPO: an open-source LLM RL system (Decoupled Clip and Dynamic sAmPling Optimization) — cited as [86], 2025-ish https://scholar.google.com/scholar?q=DAPO:+an+open-source+LLM+RL+system+(Decoupled+Clip+and+Dynamic+sAmPling+Optimization) 11. ESM2 / Language models of protein sequences at the scale of evolution — Lin et al., 2023 https://scholar.google.com/scholar?q=ESM2+/+Language+models+of+protein+sequences+at+the+scale+of+evolution Interactive Visualization: Deep Native Structural Reasoning for Proteins, Molecules, and Crystals

  8. 2d ago

    Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs

    This episode explores whether GPU dominance in AI computing could be challenged by matrix-enhanced CPUs, examining a paper by Jack Dongarra, Torsten Hoefler, and Satoshi Matsuoka from the University of Tennessee, ETH Zurich, and RIKEN. It traces why GPUs became essential for AI—citing AlexNet's 2012 breakthrough, NVIDIA's introduction of high-bandwidth memory with the P100, and tensor cores in Volta—before unpacking the two architectural bets the paper makes: on-package HBM (which physically stacks DRAM dies for a 1024-bit-wide interface versus 64 bits on conventional memory) and CPU-integrated matrix engines like ARM's SME or Intel's AMX combined with mixed-precision arithmetic. The discussion highlights Fugaku's A64FX chip as real-world proof that the bandwidth side of this equation already works, having topped the Top500 and memory-bound benchmark lists from 2020 to 2022, while noting the matrix-engine half remains a projection the paper tests on a trillion-parameter Kimi-K2 model at 256K-token context. Listeners interested in AI hardware economics will find this compelling for its rare rigor: the hosts stress that the authors clearly separate measured hardware results from projected estimates rather than blending speculation with data. The episode also breaks down the prefill-versus-decode distinction in LLM inference—compute-bound versus memory-bandwidth-bound—as the key lens for understanding where CPU architecture could realistically compete with GPUs. Sources: 1. Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs https://podcast.do-not-panic.com/uploaded-pdfs/2026-07-09T15-03-00-780Z-need-gpus.pdf 2. A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV — D. U. Lee, K. W. Kim, K. W. Kim, et al. (SK Hynix), 2014 https://scholar.google.com/scholar?q=A+1.2V+8Gb+8-Channel+128GB/s+High-Bandwidth+Memory+(HBM)+Stacked+DRAM+with+Effective+Microbump+I/O+Test+Methods+Using+29nm+Process+and+TSV 3. Co-Design for A64FX Manycore Processor and 'Fugaku' — Mitsuhisa Sato, Yutaka Ishikawa, Hirokazu Tomita, et al. (RIKEN/Fujitsu), 2020 https://scholar.google.com/scholar?q=Co-Design+for+A64FX+Manycore+Processor+and+'Fugaku' 4. Roofline: An Insightful Visual Performance Model for Multicore Architectures — Samuel Williams, Andrew Waterman, David Patterson, 2009 https://scholar.google.com/scholar?q=Roofline:+An+Insightful+Visual+Performance+Model+for+Multicore+Architectures 5. Bandwidth-Optimized Sapphire Rapids HBM (Xeon CPU Max Series) Technical Overview — Intel Corporation architecture team, 2023 https://scholar.google.com/scholar?q=Bandwidth-Optimized+Sapphire+Rapids+HBM+(Xeon+CPU+Max+Series)+Technical+Overview 6. DeepSeek-V3 Technical Report (and DeepSeek-V3.2 sparse-attention follow-up) — DeepSeek-AI, 2024-2025 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report+(and+DeepSeek-V3.2+sparse-attention+follow-up) 7. Kimi K2 Technical Report — Moonshot AI / Kimi Team, 2025 https://scholar.google.com/scholar?q=Kimi+K2+Technical+Report 8. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Y. Li, F. Wei, C. Zhang, H. Zhang (EAGLE line of work), 2024-2025 https://scholar.google.com/scholar?q=EAGLE-3:+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test 9. In-Datacenter Performance Analysis of a Tensor Processing Unit — N. Jouppi et al., 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 10. NVIDIA Grace Hopper / Grace Blackwell Superchip architecture whitepapers — NVIDIA, 2022-2025 https://scholar.google.com/scholar?q=NVIDIA+Grace+Hopper+/+Grace+Blackwell+Superchip+architecture+whitepapers Interactive Visualization: Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs

Ratings & Reviews

3.7
out of 5
3 Ratings

About

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

You Might Also Like