AI Post Transformers

mcgrof

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

  1. 2일 전

    Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents

    This episode explores "Lifelong Learning of Large Language Model based Agents: A Roadmap," a survey examining how AI agents can continuously adapt to changing environments without losing prior knowledge. The discussion centers on the stability-plasticity dilemma—the tension between preserving learned capabilities and remaining flexible enough to absorb new information—and how this classical problem from connectionist neuroscience resurfaces in a new form for modern agents that rarely fine-tune their underlying weights. Key arguments include the concept of "functional forgetting," where information technically persists in vector stores but becomes practically inaccessible if retrieval or context limits fail to surface it, and a four-part memory taxonomy spanning working, episodic, semantic, and parametric memory. The hosts also trace how this survey synthesizes and extends two separate research lineages—internal-knowledge-focused LLM surveys and agent-architecture surveys—into a unified framework modeled as a goal-conditioned POMDP. Listeners interested in why coding assistants, web-browsing agents, and other AI tools degrade over time as their environments shift will find concrete framing for that problem here. Sources: 1. Lifelong Learning of Large Language Model based Agents: A Roadmap — Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, Qianli Ma, 2025 http://arxiv.org/abs/2501.07278 2. Overcoming Catastrophic Forgetting in Neural Networks — James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, et al. (DeepMind), 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks 3. Continual Lifelong Learning with Neural Networks: A Review — German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter, 2019 https://scholar.google.com/scholar?q=Continual+Lifelong+Learning+with+Neural+Networks%3A+A+Review 4. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein (Stanford / Google), 2023 https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior 5. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar (NVIDIA, Caltech, UT Austin), 2023 https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models 6. Towards Lifelong Learning of Large Language Models: A Survey — J. Zheng, S. Qiu, C. Shi, Q. Ma, 2024 https://scholar.google.com/scholar?q=Towards+Lifelong+Learning+of+Large+Language+Models%3A+A+Survey 7. Loss of Plasticity in Deep Continual Learning — S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, R. S. Sutton, 2024 https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning 8. A Survey on Large Language Model Based Autonomous Agents — L. Wang, C. Ma, X. Feng, et al., 2024 https://scholar.google.com/scholar?q=A+Survey+on+Large+Language+Model+Based+Autonomous+Agents 9. WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models — P. Wang, Z. Li, N. Zhang, Z. Xu, Y. Yao, Y. Jiang, P. Xie, F. Huang, H. Chen, 2024 https://scholar.google.com/scholar?q=WISE%3A+Rethinking+the+Knowledge+Memory+for+Lifelong+Model+Editing+of+Large+Language+Models 10. Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem — M. McCloskey, N. J. Cohen, 1989 https://scholar.google.com/scholar?q=Catastrophic+Interference+in+Connectionist+Networks%3A+The+Sequential+Learning+Problem Interactive Visualization: Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents

  2. 2일 전

    Data Temporality's Hidden Impact on LLM Pretraining

    This episode explores why open-weight LLMs like Llama 3.1, Gemma3, Qwen3, and Olmo3 systematically lose 11–39% relative accuracy on facts from 2023–2024 compared to facts from 2020–2021, even though the more recent data falls within their training window. Drawing on Kyutai's paper "Understanding Data Temporality Impact on Large Language Models Pre-training," the discussion traces this "knowledge horizon gap" to a design choice baked into standard pretraining: corpora from many years are pooled and globally shuffled before training, erasing any timestamp signal and letting older, more frequently re-crawled data dominate. The hosts connect this to learning-rate decay schedules, arguing that data seen late in training — when updates are small and durable — gets imprinted far more strongly than data seen early, so chronological ordering (feeding snapshots 2018 through 2025 in sequence) could exploit that same mechanism to anchor recent facts instead of losing them. They situate the work against Zhao et al.'s "Set the Clock" research and Bengio's foundational curriculum-learning ideas, framing chronological training as a strikingly cheap intervention — same tokens, same compute, same architecture — for a problem the field has largely ignored. It's a compelling listen for anyone puzzling over why "knowledge cutoff" claims don't match what models actually seem to know. Sources: 1. Understanding Data Temporality Impact on Large Language Models Pre-training — Hippolyte Pilchen, Romain Fabre, Franck Signe Talla, Patrick Perez, Edouard Grave, 2026 http://arxiv.org/abs/2605.22769 2. Set the Clock: Temporal Alignment of Pretrained Language Models — Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, 2024 https://scholar.google.com/scholar?q=Set+the+Clock%3A+Temporal+Alignment+of+Pretrained+Language+Models 3. Time-Aware Language Models as Temporal Knowledge Bases — Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen, 2022 https://scholar.google.com/scholar?q=Time-Aware+Language+Models+as+Temporal+Knowledge+Bases 4. TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models — Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo, 2022 https://scholar.google.com/scholar?q=TemporalWiki%3A+A+Lifelong+Benchmark+for+Training+and+Evaluating+Ever-Evolving+Language+Models 5. RealTime QA: What's the Answer Right Now? — Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Yutaro Yamada, Deqing Fu, Tushar Khot, Ashish Sabharwal, Rik Koncel-Kedziorski, Yejin Choi, Noah A. Smith, Kentaro Inui, 2022 https://scholar.google.com/scholar?q=RealTime+QA%3A+What%27s+the+Answer+Right+Now%3F 6. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009 https://scholar.google.com/scholar?q=Curriculum+Learning 7. TimeLMs: Diachronic Language Models from Twitter — Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, Jose Camacho-Collados, 2022 https://scholar.google.com/scholar?q=TimeLMs%3A+Diachronic+Language+Models+from+Twitter 8. In-Context Pretraining: Language Modeling Beyond Document Boundaries — Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Wen-tau Yih, Mike Lewis, 2023 https://scholar.google.com/scholar?q=In-Context+Pretraining%3A+Language+Modeling+Beyond+Document+Boundaries 9. Towards Continual Knowledge Learning of Language Models — Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo, 2022 https://scholar.google.com/scholar?q=Towards+Continual+Knowledge+Learning+of+Language+Models 10. TiC-LM: A web-scale benchmark for time-continual LLM pretraining — Li, J., Armandpour, M., Mirzadeh, I., Mehta, S., Shankar, V., Vemulapalli, R., Bengio, S., Tuzel, O., Farajtabar, M., Pouransari, H., Faghri, F., 2025 https://scholar.google.com/scholar?q=TiC-LM%3A+A+web-scale+benchmark+for+time-continual+LLM+pretraining 11. How do language models learn facts? Dynamics, curricula and hallucinations — Zucchet, N., Bornschein, J., Chan, S. C., Lampinen, A. K., Pascanu, R., De, S., 2025 https://scholar.google.com/scholar?q=How+do+language+models+learn+facts%3F+Dynamics%2C+curricula+and+hallucinations 12. Data mixing can induce phase transitions in knowledge acquisition — Gu, X., Lyu, K., Li, J., Zhang, J., 2026 https://scholar.google.com/scholar?q=Data+mixing+can+induce+phase+transitions+in+knowledge+acquisition 13. TiMoE: Time-aware mixture of language experts — Faro, R., Fan, D., Alphaidze, T., Jaggi, M., 2025 https://scholar.google.com/scholar?q=TiMoE%3A+Time-aware+mixture+of+language+experts 14. Does your data spark joy? Performance gains from domain upsampling at the end of training — Blakeney, C., Paul, M., Larsen, B. W., Owen, S., Frankle, J., 2024 https://scholar.google.com/scholar?q=Does+your+data+spark+joy%3F+Performance+gains+from+domain+upsampling+at+the+end+of+training Interactive Visualization: Data Temporality's Hidden Impact on LLM Pretraining

  3. 2일 전

    Distributed Weight Data Parallelism Cuts LLM Inference Stalls

    This episode explores DWDP (Distributed Weight Data Parallelism), a new NVIDIA-authored approach to LLM inference on NVL72 systems that targets a subtle but costly inefficiency: GPUs sitting idle while they wait to synchronize with slower peers. The hosts unpack how existing model-parallelism strategies—expert, tensor, and pipeline parallelism—all share a hidden flaw, forcing every GPU to hit a synchronization barrier at each layer boundary, which the paper's own baseline shows can waste around twelve percent of total inference time even under ordinary workload imbalance. They explain why smarter scheduling alone (cache-aware or load-aware routing) can't fix this, since it only shrinks the imbalance feeding into the wait rather than eliminating the wait itself. The discussion then turns to DWDP's core idea: keeping GPUs fully data-parallel while having each one asynchronously prefetch missing expert weights from peers on demand, timed to hide the fetch behind ongoing compute. Listeners interested in the mechanics of large-scale MoE inference, GPU synchronization bottlenecks, and practical systems-level solutions to straggler problems will find the technical walkthrough especially rewarding. Sources: 1. DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72 — Wanqian Li, Jintao Peng, Zongfei Jing, Tianyu Zhang, Ze Long, Xianjie Qiao, Xiaoming Chen, Dongxu Yang, Kefeng Duan, June Yang, 2026 http://arxiv.org/abs/2604.01621 2. DeepSeek-V3 Technical Report — DeepSeek-AI, Aixin Liu, Bei Feng, et al., 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 3. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, et al., 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot 4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, et al., 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 5. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, et al., 2024 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 6. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, et al., 2023 https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale 7. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale Interactive Visualization: Distributed Weight Data Parallelism Cuts LLM Inference Stalls

  4. 3일 전

    AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC

    This episode explores AdaJEPA, an adaptive latent world model that challenges the standard "train once, freeze forever" assumption behind robot planning systems. The hosts trace the technical lineage from Yann LeCun's Joint-Embedding Predictive Architecture concept through model predictive control's decades-old roots in process engineering and rocket landing, showing how these pieces combine to let a deployed robot keep updating its internal model using only the consequences of its own actions — no new labels, demonstrations, or retraining pipeline required. Central to the discussion is how distribution shift causes small prediction errors to compound across multi-step planning horizons, and how test-time adaptation, borrowed from image classification and paralleled to cerebellar motor learning, closes that loop by treating each observed transition as a live training example. The conversation grounds abstract control theory in concrete deployment scenarios, from unfamiliar object shapes to shifting friction and lighting. Listeners interested in robotics, control theory, or self-supervised learning will find a clear walkthrough of why frozen world models fail in the wild and what it means for a model to keep learning after "training" officially ends. Sources: 1. AdaJEPA: An Adaptive Latent World Model — Ying Wang, Oumayma Bounou, Yann LeCun, Mengye Ren, 2026 http://arxiv.org/abs/2606.32026 2. Model Predictive Control: Theory and Practice — A Survey — Carlos E. García, David M. Prett, Manfred Morari, 1989 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey 3. Model Predictive Control: Classical, Robust and Stochastic — Basil Kouvaritakis, Mark Cannon, 2016 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Classical%2C+Robust+and+Stochastic 4. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) — Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine, 2018 https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+in+a+Handful+of+Trials+using+Probabilistic+Dynamics+Models+%28PETS%29 5. Learning Latent Dynamics for Planning from Pixels (PlaNet) — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2019 https://scholar.google.com/scholar?q=Learning+Latent+Dynamics+for+Planning+from+Pixels+%28PlaNet%29 6. Dino-wm: World models on pre-trained visual features enable zero-shot planning — Zhou, G., Pan, H., LeCun, Y., and Pinto, L., 2025 https://scholar.google.com/scholar?q=Dino-wm%3A+World+models+on+pre-trained+visual+features+enable+zero-shot+planning 7. Temporal straightening for latent planning — Wang, Y., Bounou, O., Zhou, G., Balestriero, R., Rudner, T. G., LeCun, Y., and Ren, M., 2026 https://scholar.google.com/scholar?q=Temporal+straightening+for+latent+planning 8. Closing the train-test gap in world models for gradient-based planning — Parthasarathy, A., Kalra, N., Agrawal, R., LeCun, Y., Bounou, O., Izmailov, P., and Goldblum, M., 2025 https://scholar.google.com/scholar?q=Closing+the+train-test+gap+in+world+models+for+gradient-based+planning 9. Td-mpc2: Scalable, robust world models for continuous control — Hansen, N., Su, H., and Wang, X., 2024 https://scholar.google.com/scholar?q=Td-mpc2%3A+Scalable%2C+robust+world+models+for+continuous+control 10. Adawm: Adaptive world model based planning for autonomous driving — Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., and Zhang, J., 2025 https://scholar.google.com/scholar?q=Adawm%3A+Adaptive+world+model+based+planning+for+autonomous+driving 11. Test-time training with self-supervision for generalization under distribution shifts — Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M., 2020 https://scholar.google.com/scholar?q=Test-time+training+with+self-supervision+for+generalization+under+distribution+shifts Interactive Visualization: AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC

  5. 3일 전

    Cross-Family Speculative Prefill Cuts Long-Context Latency

    This episode explores cross-family speculative prefill, a technique for cutting long-context inference latency by using a small "draft" model to identify which parts of a lengthy prompt matter before a much larger target model processes it. The hosts unpack why this is a hard problem in principle — draft and target models often use completely different tokenizers and architectures, meaning attention-based importance signals shouldn't obviously transfer between them — and trace the lineage from speculative decoding through the original same-family Speculative Prefill work to this paper's cross-family generalization. They highlight the practical motivation: models like DeepSeek and Kimi-K2 have no smaller sibling in their own family, so a technique that only works with matched draft/target pairs is a dead end for real deployments. Key results discussed include an 18x reduction in time-to-first-token, and the episode weighs supporting evidence from prior work on attention sinks against the stronger, less obvious claim that a full salience ranking over a 100,000-token document can transfer across unrelated architectures. Listeners interested in practical LLM efficiency techniques and the mechanics of long-context inference will find the back-and-forth skepticism over whether the method should even work, given the tokenizer mismatch, particularly engaging. Sources: 1. Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models — Shubhangi Upasani, Ravi Shanker Raju, Bo Li, Mengmeng Ji, John Long, Chen Wu, Urmish Thakker, Guangtao Wang, 2026 http://arxiv.org/abs/2603.02631 2. Speculative Prefill — Liu et al., 2025 https://scholar.google.com/scholar?q=Speculative+Prefill 3. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023 (ICLR 2024) https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 5. SnapKV: LLM Knows What You Are Looking For Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024 https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+Are+Looking+For+Before+Generation 6. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models — Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2023 (EMNLP 2023) https://scholar.google.com/scholar?q=LLMLingua%3A+Compressing+Prompts+for+Accelerated+Inference+of+Large+Language+Models 7. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression — Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2024 (ACL 2024) https://scholar.google.com/scholar?q=LongLLMLingua%3A+Accelerating+and+Enhancing+LLMs+in+Long+Context+Scenarios+via+Prompt+Compression 8. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Ruhle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, Dongmei Zhang, 2024 (ACL Findings 2024) https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression 9. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023 (NeurIPS 2023) https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens 10. Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation — Jingyu Liu, Beidi Chen, Ce Zhang, 2025 https://scholar.google.com/scholar?q=Speculative+Prefill%3A+Turbocharging+TTFT+with+Lightweight+and+Training-Free+Token+Importance+Estimation 11. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, et al. (SambaNova), 2025 https://scholar.google.com/scholar?q=SnapStream%3A+Efficient+Long+Sequence+Decoding+on+Dataflow+Accelerators 12. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — Guangtao Wang, Shubhangi Upasani, Chen Wu, et al. (SambaNova), 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference 13. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, et al., 2024 https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention Interactive Visualization: Cross-Family Speculative Prefill Cuts Long-Context Latency

  6. 3일 전

    Test-Time Training Turns EDA Feedback Into Live Weight Updates for RTL

    This episode explores Alpha-RTL, a framework applying test-time training to RTL hardware optimization, where an LLM updates its own weights live for each chip design using real EDA toolchain feedback rather than a static, pre-trained policy. The discussion contrasts this approach with two existing camps: agentic search methods (like REvolution) that iterate over a frozen model and discard synthesis feedback after each run, and training-time reinforcement learning (like ChipSeek) that learns once offline and only samples at inference. It unpacks why functional correctness in Verilog is a weak proxy for what chip teams actually optimize — PPA, the area-delay-power product measured only after synthesis — and traces the paper's core techniques back to their origins: test-time training from Sun et al.'s 2020 UC Berkeley work, and PUCT search from Kocsis and Szepesvári's 2006 UCT paper, extended here into a persistent state pool of Verilog candidates refined over gradient updates rather than resampled from scratch. Listeners interested in the mechanics of closing the loop between LLM code generation and physical design constraints — and the unusual tradeoff of burning GPU-hours to fine-tune a model for a single, disposable hardware block — will find the episode's breakdown of RLVR-style staged verification (compile, simulate, synthesize) particularly useful. Sources: 1. Alpha-RTL: Test-Time Training for RTL Hardware Optimization — Peilong Zhou, Zhirong Chen, Cangyuan Li, Haoyu Gao, Kaiyan Chang, Ziming Qu, Ying Wang, 2026 http://arxiv.org/abs/2606.05253 2. Bandit based Monte-Carlo Planning — Levente Kocsis, Csaba Szepesvári, 2006 https://scholar.google.com/scholar?q=Bandit+based+Monte-Carlo+Planning 3. A Survey of Monte Carlo Tree Search Methods — Cameron Browne, Edward Powley, Daniel Whitehouse, Simon Lucas, Peter Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, Simon Colton, 2012 https://scholar.google.com/scholar?q=A+Survey+of+Monte+Carlo+Tree+Search+Methods 4. Mastering the game of Go with deep neural networks and tree search — David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, et al. (DeepMind), 2016 https://scholar.google.com/scholar?q=Mastering+the+game+of+Go+with+deep+neural+networks+and+tree+search 5. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play — David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, et al. (DeepMind), 2018 https://scholar.google.com/scholar?q=A+general+reinforcement+learning+algorithm+that+masters+chess%2C+shogi%2C+and+Go+through+self-play 6. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero) — David Silver et al., 2017/2018 https://scholar.google.com/scholar?q=Mastering+Chess+and+Shogi+by+Self-Play+with+a+General+Reinforcement+Learning+Algorithm+%28AlphaZero%29 7. A Graph Placement Methodology for Fast Chip Design (AlphaChip) — Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, et al., 2021 https://scholar.google.com/scholar?q=A+Graph+Placement+Methodology+for+Fast+Chip+Design+%28AlphaChip%29 8. Data-Driven Offline Optimization for Architecting Hardware Accelerators (PRIME) — Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, Sergey Levine, 2021/2022 https://scholar.google.com/scholar?q=Data-Driven+Offline+Optimization+for+Architecting+Hardware+Accelerators+%28PRIME%29 9. SymbiYosys / eqy (formal equivalence checking for Yosys-based flows) — YosysHQ / Claire Wolf and contributors, ongoing https://scholar.google.com/scholar?q=SymbiYosys+%2F+eqy+%28formal+equivalence+checking+for+Yosys-based+flows%29

  7. 4일 전

    Main Trust Issue in FPGA HLS Design Workflow

    This episode examines ContractHIL-HLS, a paper from Jingbo Zhang and colleagues at Beijing University of Technology (posted to arXiv July 28, 2026) that tackles high-level synthesis for FPGA design, where LLM-generated hardware can compile cleanly and pass simulation yet still fail on real silicon due to timing violations, routing congestion, or power overruns that only surface during actual synthesis and place-and-route. The discussion contrasts this work with prior efforts like Chip-Chat, RTLLM, and HLS-Eval, arguing those prove models can generate hardware code but not that a workflow can reliably preserve design intent and incorporate tool feedback across multiple steps. Rather than relying on conversational role-prompting, where constraints can silently drift or vanish between turns, the paper's architecture splits agents by transformation type — a Contract Agent converts natural language into a structured object with named fields for interface, constraints, and validation policy, an HTML Agent renders it stably, and a Hardware-in-the-Loop Agent implements and revises designs using real Vitis HLS synthesis, Vivado place-and-route, and board bring-up rather than trusting the model's own claims. The hosts debate whether structured fields actually prevent drift better than conversational memory does, landing on the distinction that a missing field is inspectable while conversational drift is not, though enforcement remains an open question. Listeners interested in how hardware-design automation might borrow validation rigor from aerospace and control-systems engineering will find the explanation of Hardware-in-the-Loop testing, and its adaptation to catch AI-generated designs before they reach costly physical fabrication, especially compelling. Sources: 1. ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design — Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang, 2026 http://arxiv.org/abs/2607.25283 2. LegUp: High-Level Synthesis for FPGA-Based Processor/Accelerator Systems — Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Anderson, Stephen Brown, Tomasz Czajkowski, 2011 (FPGA conference; extended in ACM TODAES 2013) https://scholar.google.com/scholar?q=LegUp%3A+High-Level+Synthesis+for+FPGA-Based+Processor%2FAccelerator+Systems 3. Fast Inference of Deep Neural Networks in FPGAs for Particle Physics (hls4ml) — Javier Duarte, Song Han, Philip Harris, et al., 2018 https://scholar.google.com/scholar?q=Fast+Inference+of+Deep+Neural+Networks+in+FPGAs+for+Particle+Physics+%28hls4ml%29 4. Chip-Chat: Challenges and Opportunities in Conversational Hardware Design — Jason Blocklove, Siddharth Garg, Ramesh Karri, Hammond Pearce, 2023 https://scholar.google.com/scholar?q=Chip-Chat%3A+Challenges+and+Opportunities+in+Conversational+Hardware+Design 5. AutoChip: Automating HDL Generation Using LLM Feedback — Shailja Thakur et al., 2023 https://scholar.google.com/scholar?q=AutoChip%3A+Automating+HDL+Generation+Using+LLM+Feedback 6. MnasNet: Platform-Aware Neural Architecture Search for Mobile — Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, Quoc V. Le, 2019 https://scholar.google.com/scholar?q=MnasNet%3A+Platform-Aware+Neural+Architecture+Search+for+Mobile 7. FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search — Bichen Wu, Xiaoliang Zhang, Kaiwen Weng, Yandong Guo, Peizhao Zhang, Yanghan Wang, Kurt Keutzer, Peter Vajda, 2019 https://scholar.google.com/scholar?q=FBNet%3A+Hardware-Aware+Efficient+ConvNet+Design+via+Differentiable+Neural+Architecture+Search 8. Learning Dexterous In-Hand Manipulation — OpenAI (Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, et al.), 2020 (IJRR; earlier preprint 2018) https://scholar.google.com/scholar?q=Learning+Dexterous+In-Hand+Manipulation 9. CRYSTALS-Kyber: A CCA-Secure Module-Lattice-Based KEM — Joppe Bos, Léo Ducas, Eike Kiltz, Tancrède Lepoint, Vadim Lyubashevsky, John M. Schanck, Peter Schwabe, Gregor Seiler, Damien Stehlé, 2018 https://scholar.google.com/scholar?q=CRYSTALS-Kyber%3A+A+CCA-Secure+Module-Lattice-Based+KEM 10. A Compact Hardware Implementation of CCA-Secure Key Exchange Mechanism CRYSTALS-KYBER on FPGA — Yufei Xing, Shuguo Li, 2021 https://scholar.google.com/scholar?q=A+Compact+Hardware+Implementation+of+CCA-Secure+Key+Exchange+Mechanism+CRYSTALS-KYBER+on+FPGA 11. Module-Lattice-Based Key-Encapsulation Mechanism Standard (FIPS 203) — National Institute of Standards and Technology (NIST), 2024 https://scholar.google.com/scholar?q=Module-Lattice-Based+Key-Encapsulation+Mechanism+Standard+%28FIPS+203%29 12. KyberMat and CRYPHTOR (accelerator designs cited directly in the ContractHIL-HLS paper) — Not independently verified here — cited by the ContractHIL-HLS authors as references [8] and [9], Recent (post-2023, exact years unconfirmed) https://scholar.google.com/scholar?q=KyberMat+and+CRYPHTOR+%28accelerator+designs+cited+directly+in+the+ContractHIL-HLS+paper%29 13. HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks — S. Abi-Karam, C. Hao, 2025 https://scholar.google.com/scholar?q=HLS-Eval%3A+A+Benchmark+and+Framework+for+Evaluating+LLMs+on+High-Level+Synthesis+Design+Tasks 14. SAGE-HLS: Syntax-Aware AST-Guided LLM for High-Level Synthesis Code Generation — M. Z. S. Khan, N. Mashnoor, M. Akyash et al., 2025 https://scholar.google.com/scholar?q=SAGE-HLS%3A+Syntax-Aware+AST-Guided+LLM+for+High-Level+Synthesis+Code+Generation 15. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT — J. White, Q. Fu, S. Hays et al., 2023 https://scholar.google.com/scholar?q=A+Prompt+Pattern+Catalog+to+Enhance+Prompt+Engineering+with+ChatGPT 16. Evaluating Large Language Models Trained on Code — M. Chen, J. Tworek, H. Jun et al. (OpenAI Codex/HumanEval), 2021 https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code 17. KyberMat: Efficient Accelerator for Matrix-Vector Polynomial Multiplication in CRYSTALS-Kyber via NTT and Polyphase Decomposition — W. Tan, Y. Lao, K. K. Parhi, 2023 https://scholar.google.com/scholar?q=KyberMat%3A+Efficient+Accelerator+for+Matrix-Vector+Polynomial+Multiplication+in+CRYSTALS-Kyber+via+NTT+and+Polyphase+Decomposition Interactive Visualization: Main Trust Issue in FPGA HLS Design Workflow

  8. 4일 전

    MemPO: Teaching Agents to Write Their Own Memory

    This episode explores MemPO, a self-memory policy optimization framework for long-horizon AI agents developed by researchers at Tsinghua University and Alibaba's Tongyi Lab. The discussion contrasts MemPO's approach against the dominant ReAct pattern, which accumulates full interaction history and suffers from both ballooning token costs and the "lost in the middle" degradation documented in prior research, as well as against passive retrieval-based memory systems like MemGPT and Mem0 that rely on embedding similarity rather than task outcomes. The hosts unpack how MemPO trains an agent to write compressed memory notes as a learned, RL-optimized action — discarding raw tool outputs and reasoning traces at each step in favor of a single distilled note — using Group Relative Policy Optimization to solve the credit-assignment problem of rewarding intermediate memory decisions from a single end-of-trajectory success signal. Listeners interested in agent architecture, RL training objectives, or the tradeoffs between context-window scaling and structured memory will find the episode's walk-through of the mem/think/tool_call decomposition particularly useful. The conversation also traces the intellectual lineage of the ideas, from Minsky's original framing of credit assignment to retrieval-augmented generation's origins at Facebook AI Research. Sources: 1. MemPO: Self-Memory Policy Optimization for Long-Horizon Agents — Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo, 2026 http://arxiv.org/abs/2603.00680 2. Policy Gradient Methods for Reinforcement Learning with Function Approximation — Richard S. Sutton, David McAllester, Satinder Singh, Yishay Mansour, 1999/2000 https://scholar.google.com/scholar?q=Policy+Gradient+Methods+for+Reinforcement+Learning+with+Function+Approximation 3. High-Dimensional Continuous Control Using Generalized Advantage Estimation — John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel, 2016 https://scholar.google.com/scholar?q=High-Dimensional+Continuous+Control+Using+Generalized+Advantage+Estimation 4. RUDDER: Return Decomposition for Delayed Rewards — Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Adler, Johannes Brandstetter, Sepp Hochreiter, 2019 https://scholar.google.com/scholar?q=RUDDER%3A+Return+Decomposition+for+Delayed+Rewards 5. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023 https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step 6. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents — Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying, 2025 https://scholar.google.com/scholar?q=Information+Gain-based+Policy+Optimization%3A+A+Simple+and+Effective+Approach+for+Multi-Turn+LLM+Agents 7. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents — Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang, 2025 https://scholar.google.com/scholar?q=MEM1%3A+Learning+to+Synergize+Memory+and+Reasoning+for+Efficient+Long-Horizon+Agents 8. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023 https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning 9. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning — Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, 2025 https://scholar.google.com/scholar?q=Search-R1%3A+Training+LLMs+to+Reason+and+Leverage+Search+Engines+with+Reinforcement+Learning 10. Attnpo: Attention-guided Process Supervision for Efficient Reasoning — Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu, 2026 https://scholar.google.com/scholar?q=Attnpo%3A+Attention-guided+Process+Supervision+for+Efficient+Reasoning 11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2024 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems Interactive Visualization: MemPO: Teaching Agents to Write Their Own Memory

소개

AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.

좋아할 만한 다른 항목