AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. há 2 dias

    Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup

    This episode digs into a paper testing whether hand-written PTX assembly beats NVIDIA's WMMA API for Tensor Core GEMM kernels on an L4 GPU, finding that the answer flips depending on numeric precision rather than holding as a universal rule. The discussion covers the hardware distinction between Tensor Cores and regular CUDA cores, and contrasts the convenience of WMMA against the finer control PTX offers through instructions like cp.async, ldmatrix, and mma.sync. A key thread traces why this matters in practice: quantized LLM serving at INT8 or INT4 shifts kernels from compute-bound to memory-bound, making the precision-dependent payoff of hand-tuned PTX directly relevant to running open-weight models cheaply. The episode also addresses the methodological choice to test on a single GPU, arguing that isolating precision and instruction-set effects requires holding hardware constant rather than spreading across devices. Listeners get a concrete framework for deciding when the extra engineering effort of writing raw PTX is worth it versus when it's wasted work. Sources: 1. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 — Matt J. Borowski, Blazej Osinski, 2026 http://arxiv.org/abs/2608.10103 2. NVIDIA Tensor Core Programmability, Performance & Precision — Stefano Markidis, Steven W. D. Chien, Erwin Laure, Ivy B. Peng, Jeffrey S. Vetter, 2018 https://scholar.google.com/scholar?q=NVIDIA+Tensor+Core+Programmability%2C+Performance+%26+Precision 3. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2018 https://scholar.google.com/scholar?q=Dissecting+the+NVIDIA+Volta+GPU+Architecture+via+Microbenchmarking 4. CUTLASS: CUDA Templates for Linear Algebra Subroutines — Andrew Kerr, Duane Merrill, Julien Demouth, John Tran (NVIDIA), with ongoing project contributors, 2018 (initial release, actively maintained since) https://scholar.google.com/scholar?q=CUTLASS%3A+CUDA+Templates+for+Linear+Algebra+Subroutines 5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 6. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019 https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations 7. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020 (OSDI) https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning 8. Dissecting the Ampere GPU Architecture via Microbenchmarking — Wei Sun, Ang Li, Tong Geng, Sander Stuijk, Henk Corporaal, 2022 https://scholar.google.com/scholar?q=Dissecting+the+Ampere+GPU+Architecture+via+Microbenchmarking 9. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 (OSDI) https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning 10. Understanding Latency Hiding on GPUs — V. Volkov, 2016 https://scholar.google.com/scholar?q=Understanding+Latency+Hiding+on+GPUs 11. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — J. Lin et al., 2024 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 12. FlashInfer: Kernel Library for LLM Serving — Z. Ye et al., 2024 https://scholar.google.com/scholar?q=FlashInfer%3A+Kernel+Library+for+LLM+Serving

  2. há 2 dias

    Inside Claude Code's Agentic Loop: A Design Space Analysis

    This episode dissects the internal architecture of Claude Code by examining its extracted TypeScript source (v2.1.88) alongside two other agent systems, OpenClaw and Hermes Agent, drawing on the paper "Dive into Claude Code" by Jiacheng Liu et al. from VILA Lab at MBZUAI. It reveals that the model's reasoning core is essentially a single while-loop — literally called queryLoop() — with everything else (permissions, context management, tools, subagents) built as scaffolding around it. The discussion covers deny-first permission rules, the five-stage compaction pipeline for managing context windows, subagent delegation with isolated context windows, and how the Model Context Protocol connects to external tool servers. The hosts also extract five human values embedded directly in the code — human decision authority, safety/security/privacy, reliable execution, capability amplification, and contextual adaptability — framing the episode as less about how Claude Code works and more about what its designers chose to prioritize, made visible through actual implementation choices. Sources: 1. Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems — Jiacheng Liu, Xiaohan Zhao, Xinyi Shang, Zhiqiang Shen, 2026 http://arxiv.org/abs/2604.14228 2. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Mario Fritz, 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 3. GAIA: A Benchmark for General AI Assistants — Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom, 2023 https://scholar.google.com/scholar?q=GAIA%3A+A+Benchmark+for+General+AI+Assistants 4. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al., 2022 https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback 5. AgentBench: Evaluating LLMs as Agents — Xiao Liu, Hao Yu, Hanchen Zhang, et al., 2023 https://scholar.google.com/scholar?q=AgentBench%3A+Evaluating+LLMs+as+Agents

  3. há 2 dias

    Mechanist: Automating the Discovery of How AI Models Think

    This episode examines "Mechanist," a multi-agent system built by researchers at Zhejiang University, NUS, Southern University of Science and Technology, Heriot-Watt, UC San Diego, and Northeastern University to automate mechanistic interpretability research itself, rather than automating experiments in an external domain like chemistry or biology. The discussion covers how a central orchestrator coordinates four agents—hypothesis, experiment, verification, and iteration—drawing on a 13,000-paper interpretability knowledge graph and a 43-million-paper cross-disciplinary graph called SciAtlas to generate and test theories about how models actually compute. Key concepts explored include subliminal learning, where a trait transfers from teacher to student model through data that looks unrelated to it, and a three-frame belief decomposition (World Knowledge, Personal Belief, Attributed Belief) used to probe whether models genuinely separate fact from attributed belief. The episode previews four escalating case studies, starting with the discovery of a previously unflagged multimodal safety risk and building toward using mechanistic theories to directly intervene on model internals and even steer a biological system, raising the question of whether an AI system can meaningfully explain the black box that produced it. Sources: 1. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen, 2026 http://arxiv.org/abs/2608.12036 2. Language models transmit behavioural traits through hidden signals in data — Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Sören Mindermann, Jacob Hilton, Samuel Marks, Owain Evans, 2026 (Nature) https://scholar.google.com/scholar?q=Language+models+transmit+behavioural+traits+through+hidden+signals+in+data 3. Subliminal learning is a lora artifact — Todd Nief, Harvey Yiyun Fu, Mark Muchane, Ari Holtzman, 2026 https://scholar.google.com/scholar?q=Subliminal+learning+is+a+lora+artifact 4. Language models cannot reliably distinguish belief from knowledge and fact — Mirac Suzgun, Tayfun Gur, Federico Bianchi, Daniel E. Ho, Thomas Icard, Dan Jurafsky, James Zou, 2025 (Nature Machine Intelligence) https://scholar.google.com/scholar?q=Language+models+cannot+reliably+distinguish+belief+from+knowledge+and+fact 5. Genome modelling and design across all domains of life with evo 2 — Garyk Brixi, Matthew G. Durrant, Jerome Ku, et al., 2026 (Nature) https://scholar.google.com/scholar?q=Genome+modelling+and+design+across+all+domains+of+life+with+evo+2 6. Towards end-to-end automation of ai research — Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune, 2026 (Nature) https://scholar.google.com/scholar?q=Towards+end-to-end+automation+of+ai+research 7. Sleeper agents: Training deceptive llms that persist through safety training — Evan Hubinger et al., 2024 https://scholar.google.com/scholar?q=Sleeper+agents%3A+Training+deceptive+llms+that+persist+through+safety+training 8. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers — Jan Dubinski, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans, 2026 https://scholar.google.com/scholar?q=Conditional+misalignment%3A+common+interventions+can+hide+emergent+misalignment+behind+contextual+triggers

  4. há 1 dia

    Second-Order Optimization Meets Runtime Scheduling at Scale

    This episode examines "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," a May 2026 Oxford paper introducing Asteria, a runtime system rather than a new optimizer. The discussion covers why curvature-aware methods like Shampoo and SOAP have never displaced AdamW despite converging in fewer steps, tracing the lineage from K-FAC through Distributed Shampoo to SOAP and explaining the Kronecker-factorization tricks that make tracking curvature tractable at all. The hosts unpack the paper's "three physical walls" framework — a vertical capacity wall from single-GPU memory limits, an overlap disruption wall where cubic-cost matrix operations stall compute-communication overlap, and a global consensus wall from synchronous full-state updates across mismatched network speeds — and debate whether reengineering the plumbing around an unchanged optimizer counts as a genuine research contribution. Listeners interested in distributed training infrastructure, optimizer design trade-offs, or the gap between algorithmic elegance and practical deployability will find the back-and-forth over real benchmark numbers (96 seconds versus 1.5 seconds per step) especially grounded. Sources: 1. Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training — Yishun Lu, Junhao Zhang, Zeyu Yang, Wes Armour, 2026 http://arxiv.org/abs/2605.16184 2. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 6. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — H.-J. M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat, 2023 https://scholar.google.com/scholar?q=A+Distributed+Data-Parallel+PyTorch+Implementation+of+the+Distributed+Shampoo+Optimizer+for+Training+Neural+Networks+At-Scale 7. Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading — A. Maurya, J. Ye, M. M. Rafique, F. Cappello, B. Nicolae, 2024 https://scholar.google.com/scholar?q=Deep+Optimizer+States%3A+Towards+Scalable+Training+of+Transformer+Models+Using+Interleaved+Offloading 8. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, Y. He, 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 9. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization — W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, R. B. Grosse, 2026 https://scholar.google.com/scholar?q=Understanding+and+Improving+Shampoo+and+SOAP+via+Kullback-Leibler+Minimization 10. Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training — Y. Lu, W. Armour, 2026 https://scholar.google.com/scholar?q=Beyond+the+Mean%3A+Fisher-Orthogonal+Projection+for+Natural+Gradient+Descent+in+Large+Batch+Training Interactive Visualization: Second-Order Optimization Meets Runtime Scheduling at Scale

  5. há 4 dias

    Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

    This episode examines "Is Grep All You Need? How Agent Harnesses Reshape Agentic Search," which tests whether simple regex-based retrieval can outperform vector search inside agentic pipelines like Chronos, Claude Code, and Codex CLI. The hosts dig into how retrieval mode interacts with harness architecture, delivery method (inline vs. programmatic), and backbone model choice, finding that inline grep beats inline vector search across every harness-model pairing tested — with gaps as wide as twenty points and swings as large as switching harnesses entirely. A striking case shows the same model scoring 93.1% on one harness but only 76.7% on another, suggesting orchestration and prompt construction matter as much as the retrieval algorithm itself. The discussion also surfaces a counterintuitive twist: forcing an agent to read retrieved results from a file instead of getting them dumped inline can nearly halve accuracy, even with identical underlying search. Listeners interested in RAG, agent design, or LLM evaluation methodology will find the paper's tangled-but-honest approach to measuring real deployed systems a useful corrective to cleaner but less realistic ablation studies. Sources: 1. Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sahil Sen, Akhil Kasturi, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 http://arxiv.org/abs/2605.15184 2. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022 https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models 3. WebGPT: Browser-assisted question-answering with human feedback — Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al. (OpenAI), 2021 https://scholar.google.com/scholar?q=WebGPT%3A+Browser-assisted+question-answering+with+human+feedback 4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks 5. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Di Wu, Hongwei Wang, Wenhao Yu, et al., 2024 https://scholar.google.com/scholar?q=LongMemEval%3A+Benchmarking+Chat+Assistants+on+Long-Term+Interactive+Memory 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond 7. Dense Passage Retrieval for Open-Domain Question Answering — Vladimir Karpukhin, Barlas Oğuz, Sewon Min, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Dense+Passage+Retrieval+for+Open-Domain+Question+Answering 8. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, et al., 2023 https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts 9. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Thibault Formal, Benjamin Piwowarski, Stéphane Clinchant, 2021 https://scholar.google.com/scholar?q=SPLADE%3A+Sparse+Lexical+and+Expansion+Model+for+First+Stage+Ranking 10. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models — Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych, 2021 https://scholar.google.com/scholar?q=BEIR%3A+A+Heterogenous+Benchmark+for+Zero-shot+Evaluation+of+Information+Retrieval+Models 11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 12. Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory — Sahil Sen, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 https://scholar.google.com/scholar?q=Chronos%3A+Temporal-Aware+Conversational+Agents+with+Structured+Event+Retrieval+for+Long-Term+Memory Interactive Visualization: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

  6. há 5 dias

    AI-Generated Text and the Death of the Open Web

    This episode examines a study analyzing the growth of AI-generated content across the open web, drawing on 33 monthly samples from the Internet Archive's Wayback Machine between August 2022 and May 2025. It highlights the paper's central finding that AI-generated or AI-assisted content on newly published websites rose from zero before ChatGPT's launch to roughly 35 percent by mid-2025, and explores how the authors transform "Dead Internet Theory" from internet folklore into six testable hypotheses — including semantic contraction, truth decay, positivity shift, epistemic islands, entropy dilution, and stylistic monoculture. The discussion covers the methodology behind sampling a representative slice of the internet, including logarithmic downsampling and stratification across time, MIME type, and domain to avoid bias toward heavily-crawled sites. It also connects the findings to the concept of model collapse, framing the 35 percent figure as empirical evidence for a previously theoretical concern about AI models training on their own synthetic output. Listeners interested in web ecosystem health, LLM training data quality, or the intersection of internet culture and rigorous data science will find the episode's blend of meme-to-metric translation particularly compelling. Sources: 1. The Impact of AI-Generated Text on the Internet — Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek, 2026 http://arxiv.org/abs/2604.26965 2. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature — Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn, 2023 https://scholar.google.com/scholar?q=DetectGPT%3A+Zero-Shot+Machine-Generated+Text+Detection+using+Probability+Curvature 3. Can AI-Generated Text be Reliably Detected? — Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, Soheil Feizi, 2023 https://scholar.google.com/scholar?q=Can+AI-Generated+Text+be+Reliably+Detected%3F 4. A Watermark for Large Language Models — John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023 https://scholar.google.com/scholar?q=A+Watermark+for+Large+Language+Models 5. GLTR: Statistical Detection and Visualization of Generated Text — Sebastian Gehrmann, Hendrik Strobelt, Alexander M. Rush, 2019 https://scholar.google.com/scholar?q=GLTR%3A+Statistical+Detection+and+Visualization+of+Generated+Text 6. The Curse of Recursion: Training on Generated Data Makes Models Forget — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, Ross Anderson, 2024 (Nature, updated from 2023 preprint) https://scholar.google.com/scholar?q=The+Curse+of+Recursion%3A+Training+on+Generated+Data+Makes+Models+Forget 7. Is the Internet Dead? Evaluating Claims about Web Homogenization from AI-Generated Content (representative of the broader 2024-2025 measurement literature, e.g. Muzumdar et al. and related environmental-scanning studies of Dead Internet Theory discourse) — Various (this cluster of work is cited in the paper as Muzumdar et al., 2025 and similar), 2025 https://scholar.google.com/scholar?q=Is+the+Internet+Dead%3F+Evaluating+Claims+about+Web+Homogenization+from+AI-Generated+Content+%28representative+of+the+broader+2024-2025+measurement+literature%2C+e.g.+Muzumdar+et+al.+and+related+environmental-scanning+studies+of+Dead+Internet+Theory+discourse%29 8. Studies on bot/inauthentic-content prevalence on specific platforms (e.g., La Cava et al. 2025 on social media, Matatov et al. 2024 on platform-specific AI content) — Lucio La Cava et al.; Jonathan Matatov et al. (representative platform-specific studies cited in the paper's related-work section), 2024-2025 https://scholar.google.com/scholar?q=Studies+on+bot%2Finauthentic-content+prevalence+on+specific+platforms+%28e.g.%2C+La+Cava+et+al.+2025+on+social+media%2C+Matatov+et+al.+2024+on+platform-specific+AI+content%29 9. Public Trust and Perceptions of Artificial Intelligence (Ipsos / Reuters Institute Digital News Report and Edelman Trust Barometer AI-focused editions) — Ipsos (various); Reuters Institute for the Study of Journalism (Nic Newman et al.); Edelman Trust Barometer team, 2023-2025 (recurring annual) https://scholar.google.com/scholar?q=Public+Trust+and+Perceptions+of+Artificial+Intelligence+%28Ipsos+%2F+Reuters+Institute+Digital+News+Report+and+Edelman+Trust+Barometer+AI-focused+editions%29 10. Americans' Views of Artificial Intelligence (Pew Research Center recurring survey series) — Pew Research Center (Alec Tyson, Emma Kikuchi, and colleagues), 2023-2025 (recurring) https://scholar.google.com/scholar?q=Americans%27+Views+of+Artificial+Intelligence+%28Pew+Research+Center+recurring+survey+series%29 11. The Perception Gap: Comparing Public Beliefs about Misinformation to Empirical Prevalence Estimates (representative of the risk-perception vs. measured-prevalence literature this paper's framing descends from, e.g. work following Duffy et al. and general misperception-of-misinformation-prevalence studies) — Andrew Guess, colleagues in the misinformation-prevalence research cluster (representative of this line, distinct from the AI-specific surveys above), 2019-2023 (foundational misinformation-perception literature) https://scholar.google.com/scholar?q=The+Perception+Gap%3A+Comparing+Public+Beliefs+about+Misinformation+to+Empirical+Prevalence+Estimates+%28representative+of+the+risk-perception+vs.+measured-prevalence+literature+this+paper%27s+framing+descends+from%2C+e.g.+work+following+Duffy+et+al.+and+general+misperception-of-misinformation-prevalence+studies%29 12. Documenting the English Colossal Clean Crawled Corpus (C4) — Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner, 2021 https://scholar.google.com/scholar?q=Documenting+the+English+Colossal+Clean+Crawled+Corpus+%28C4%29 13. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only — Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay, 2023 https://scholar.google.com/scholar?q=The+RefinedWeb+Dataset+for+Falcon+LLM%3A+Outperforming+Curated+Corpora+with+Web+Data%2C+and+Web+Data+Only 14. Quantifying Memorization Across Neural Language Models / broader Common Crawl representativeness and bias studies (e.g., work on Common Crawl's domain and language skew) — Various (Common Crawl bias/representativeness literature, e.g. work by Luccioni & Viviano on Common Crawl content quality, and follow-on studies), 2021-2023 https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models+%2F+broader+Common+Crawl+representativeness+and+bias+studies+%28e.g.%2C+work+on+Common+Crawl%27s+domain+and+language+skew%29 15. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, et al. (Allen Institute for AI), 2024 https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research 16. AI models collapse when trained on recursively generated data — I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal, 2024 https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data 17. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks — V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, 2025 https://scholar.google.com/scholar?q=Can+AI-generated+text+be+reliably+detected%3F+Stress+testing+AI+text+detectors+under+various+attacks 18. RAID: A shared benchmark for robust evaluation of machine-generated text detectors — L. Dugan, A. Hwang, F. Trhlik, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, C. Callison-Burch, 2024 https://scholar.google.com/scholar?q=RAID%3A+A+shared+benchmark+for+robust+evaluation+of+machine-generated+text+detectors 19. When incentives backfire, data stops being human — S. Santy, P. Bhattacharya, M. H. Ribeiro, K. Allen, S. Oh, 2025 https://scholar.google.com/scholar?q=When+incentives+backfire%2C+data+stops+being+human 20. Longitudinal sampling of URLs from the Wayback Machine — K. Garg, S. Alam, D. Ayala, M. Graham, M. C. Weigle, M. L. Nelson, 2025 https://scholar.google.com/scholar?q=Longitudinal+sampling+of+URLs+from+the+Wayback+Machine 21. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity — J. Zhang, S. Yu, D. Chong, A. Sicilia, M. R. Tomz, C. D. Manning, W. Shi, 2025 https://scholar.google.com/scholar?q=Verbalized+sampling%3A+How+to+mitigate+mode+collapse+and+unlock+LLM+diversity Interactive Visualization: AI-Generated Text and the Death of the Open Web

  7. há 5 dias

    Approaching Shannon Bound: Lossless LLM Weight Compression

    This episode explores "Approaching Shannon Bound with Lossless LLM Weight Compression," which argues that model weights stored in formats like bf16 carry far less real information than their bit-width implies—entropy measurements across six models and seven numeric formats show gaps of several bits per weight that can be recovered without any change to the underlying values. The discussion covers why memory capacity and bandwidth, not raw compute, are the real bottleneck in GPU inference, and why generic compressors like gzip fail on IEEE-754 floating point data. The hosts dig into Asymmetric Numeral Systems (ANS) as the key engineering breakthrough, since it decodes fast enough and in a tile-parallel enough fashion to run inside a live GPU kernel without becoming a new bottleneck itself. Listeners interested in the intersection of information theory and practical LLM serving will find the walkthrough of how lossless compression differs fundamentally from quantization methods like int4 or AWQ particularly compelling. Sources: 1. Approaching Shannon Bound with Lossless LLM Weight Compression — Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He, 2026 http://arxiv.org/abs/2606.15789 2. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float — Tianyi Zhang, Yang Sui, Shaochen (Henry) Zhong, et al., 2025 https://scholar.google.com/scholar?q=70%25+Size%2C+100%25+Accuracy%3A+Lossless+LLM+Compression+for+Efficient+GPU+Inference+via+Dynamic-Length+Float 3. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks — Yongchang Hao, Yanshuai Cao, Lili Mou, 2024 https://scholar.google.com/scholar?q=NeuZip%3A+Memory-Efficient+Training+and+Inference+with+Dynamic+Compression+of+Neural+Networks 4. ZipNN: Lossless Compression for AI Models — Moshik Hershcovitch, Andrew Wood, Leshem Choshen, et al. (IBM Research), 2024 https://scholar.google.com/scholar?q=ZipNN%3A+Lossless+Compression+for+AI+Models 5. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016 https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding 6. Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding — Jarek Duda, 2013 https://scholar.google.com/scholar?q=Asymmetric+Numeral+Systems%3A+Entropy+Coding+Combining+Speed+of+Huffman+Coding+with+Compression+Rate+of+Arithmetic+Coding 7. The Use of Asymmetric Numeral Systems as an Accurate Replacement for Huffman Coding — Jarek Duda, Khalid Tahboub, Neeraj J. Gadgil, Edward J. Delp, 2015 https://scholar.google.com/scholar?q=The+Use+of+Asymmetric+Numeral+Systems+as+an+Accurate+Replacement+for+Huffman+Coding 8. Variational Image Compression with a Scale Hyperprior — Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, Nick Johnston, 2018 https://scholar.google.com/scholar?q=Variational+Image+Compression+with+a+Scale+Hyperprior 9. Zstandard Compression and the application/zstd Media Type (RFC 8878) — Yann Collet, Murray Kucherawy (eds.), 2020 https://scholar.google.com/scholar?q=Zstandard+Compression+and+the+application%2Fzstd+Media+Type+%28RFC+8878%29 10. GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMs — Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, Tuo Zhao, 2024 https://scholar.google.com/scholar?q=GEAR%3A+An+Efficient+KV+Cache+Compression+Recipe+for+Near-Lossless+Generative+Inference+of+LLMs 11. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters 12. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024 https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving 13. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are All You Need — M. Davies, N. Crago, K. Sankaralingam, C. Kozyrakis, 2025 https://scholar.google.com/scholar?q=Efficient+LLM+Inference%3A+Bandwidth%2C+Compute%2C+Synchronization%2C+and+Capacity+are+All+You+Need Interactive Visualization: Approaching Shannon Bound: Lossless LLM Weight Compression

  8. há 5 dias

    Decomposing Speedups Across Runtime, Kernel, and Quantization

    This episode dissects a common but misleading claim in LLM serving benchmarks: that swapping to a faster stack alone explains a headline speedup number. It walks through the mechanics separating prefill's compute-bound math from decode's memory-bandwidth-bound token generation, then explains how continuous batching and PagedAttention keep GPUs saturated, and how GPTQ quantization paired with the Marlin kernel avoids costly dequantization overhead. By introducing a matched FP16 intermediate stack, the paper cleanly splits a reported speedup into a runtime factor and a kernel-plus-quantization factor that multiply back to the observed total. The headline finding is striking: a controlled, apples-to-apples comparison yields a 2.58x speedup dominated by runtime (68%), while a more common operational comparison — pitting a batch-capped baseline against a fully loaded modern stack — inflates that to 10.62x, with runtime's share climbing to 86%. Listeners get a rare, rigorous look at how much of the "free lunch" in inference speedups is real architecture versus benchmark framing. Sources: 1. Decomposing Speedups Across Runtime, Kernel, and Quantization https://arxiv.org/pdf/2607.11368 2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 3. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale 4. SqueezeLLM: Dense-and-Sparse Quantization — Sehoon Kim, Coleman Hooper, Amir Gholami, et al., 2023 https://scholar.google.com/scholar?q=SqueezeLLM%3A+Dense-and-Sparse+Quantization 5. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, et al., 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML Interactive Visualization: Decomposing Speedups Across Runtime, Kernel, and Quantization

Classificações e avaliações

3,7
de 5
3 avaliações

Sobre

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

Você também pode gostar de