HuggingFace 每日AI论文速递

duan

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

  1. 1d ago

    2026.09.25 | 世界模型物理推理可评测;大模型线性叠加双想法

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:27] 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性)[01:12] 🧠 Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs(你的Transformer能同时容纳两个想法:LLM中线性叠加的证据)[01:53] 🎬 WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation(WanPE:面向现代文本到视频生成的电影级提示增强)[02:40] 🧭 OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents(OmniEcho:面向全模态具身智能体的音视频空间理解)[03:18] 🤖 Agent-Editing World Model: Rethinking World Modeling for LLM Agents(智能体编辑世界模型:为LLM智能体重新思考世界建模)[04:05] 🧩 Parts-of-Speech as Emergent Categories in SAE Latent Space(词性作为SAE潜空间中的涌现类别)[04:46] 🤖 Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents(Qwen-Planner-Agent:面向真实世界移动规划智能体的闭环 AI-for-AI 框架)[05:20] 🔍 IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis(IterSynth:通过角色解耦的迭代合成重新思考深度搜索智能体)[06:04] 🤖 Coding Agents for Generalized Task and Motion Planning Problems(面向泛化任务与运动规划问题的编码智能体)[06:49] 🧠 Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone(神经谱容量:仅凭网络规格测量与设计架构)[07:28] 🛡 AgentKernel: The Trust-Native Agentic Operating System(AgentKernel:信任原生的智能体操作系统)[08:15] 🧪 Rufus-Air: An Open LLM Post-Training Recipe(Rufus-Air:开放的大语言模型后训练配方)[08:59] 🤖 PUBG Ally: A Conversational Embodied Agent as an AI Teammate(PUBG Ally:作为AI队友的对话式具身智能体)[09:44] 🤖 World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal(世界动作智能体:利用视觉语言模型通过世界动作预演实现机器人操作)[10:27] 🛸 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds(ExplorationBench:测量 AI 系统在可验证异星世界中的探索能力) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  2. 2d ago

    2026.09.24 | 说话者双轨记忆;交互学习空间推理

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:28] 🧠 SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue(SpeakerMem-R1:以说话者为中心的多方对话双轨记忆)[01:13] 🤖 Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World(Spatial-Interactor:通过与可观测物理世界交互学习空间推理)[01:59] 🌍 HappyWorld-Bench(快乐世界基准(HappyWorld-Bench))[02:40] 🧠 The Past Frames the Future: Memory for Autoregressive Video Generation(过往帧塑造未来:自回归视频生成中的记忆机制)[03:29] 🧠 Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents(即时记忆:学习为LLM智能体整理任务自适应记忆)[04:11] 🎬 RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling(RewardVerse:面向视频奖励建模的评分量规引导策略优化)[05:00] 🎯 PACT: From Credit Assignment to Critic Alignment(PACT:从信用分配到评论家对齐)[05:43] 🧪 Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?(薛定谔的代码仓库:LLM 是学会了 SWE-bench,还是记住了它?)[06:24] 📐 GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression(GeoPair:用于免训练 Transformer 压缩的几何保持跨层因子分解)[07:00] 📦 PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing(PackLab:在机器人装箱中开发、训练与评估多模态大语言模型的综合框架)[07:53] 🧠 MemBodied: Recurrent Associative Memory for Vision-Language-Action Models(MemBodied:面向视觉-语言-动作模型的循环联想记忆)[08:39] 🧪 WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents(WhatWorkedBench:基准测试AI智能体的实验理解能力)[09:21] 🧠 Hunyuan-A13B Technical Report(混元-A13B 技术报告)[10:03] 🎮 Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms(可验证隐藏动力学游戏:从已求解机制生成智能体强化学习环境)[10:45] 🎥 All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation(所有模态都是平等的,但视频更平等:弥合联合视频生成中的交叉注意力差距) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  3. 5d ago

    2026.09.22 | VLM智能迁移至机器人控制;隐式3D记忆构建视频世界模型

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:28] 🤖 Transferring the Intelligence of VLMs to Robotic Control(将VLM的智能迁移至机器人控制)[01:11] 🎥 WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory(WorldCrafter:具有隐式3D感知记忆的一致视频世界模型)[01:53] 🎮 GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay(GameHorizon 套件:游戏玩法中的多时间跨度数据与评估)[02:35] 🧬 RRSI: Regularized Recursive Self-Improvement of Agent Harnesses(RRSI:智能体运行框架的正则化递归自我改进)[03:20] 📄 Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion(文档检索感知分块(D-RAC):通过 PDF 规范化与多模态 Markdown 转换实现企业文档的通用检索感知摄取)[04:01] 🎓 OmniEdu: Open Foundation Models for Learning and Teaching(OmniEdu:面向学习与教学的开放基础模型)[04:47] 🎬 VideoGen-Agent: Reinforcing Video Generation Agents(VideoGen-Agent:强化视频生成智能体)[05:37] 🐼 onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction(onPanda:通过Token级纠正为LLM与智能体高效标注同策略对齐数据)[06:19] 🤖 Grounded Action Model: 3D Grounding as a Foundation for Robotics(接地动作模型:以3D接地作为机器人学基础)[07:05] 🤖 One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents(从一到多,从多到一:面向软件工程智能体的类别感知迭代专家训练)[07:57] 🎭 Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations(Deep Persona:面向角色扮演智能体与模拟的心理学基础架构与评估框架)[08:36] 🧠 Harness-Zero: Harness Distillation via Agent-as-Harness(Harness-Zero:通过智能体作为外壳进行外壳蒸馏)[09:18] 🤖 CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies(CARE:面向视觉-语言-动作策略的经验引导式原子纠正执行)[10:13] 🧠 Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents(Jev-Mem:面向高效 AI 智能体的 System-One 控制型智能体记忆)[10:52] 🎥 Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms(为什么视频扩散模型会违背物理规律?揭示注意力机制中的缺陷) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  4. 6d ago

    2026.09.21 | 代码合成有据技能;源码扩展编程RL环境

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:30] 🧩 Grounded Skill Synthesis from Code at Scale for Agentic Intelligence(面向智能体智能的从大规模代码中合成有依据技能)[01:21] 🤖 CodeMidas: Scaling Agentic Coding RL Environments from Code Itself(CodeMidas:从代码本身扩展智能体编程强化学习环境)[02:03] 🧬 EvoOntology: A Self-Evolving Ontology Layer for Data Agents(EvoOntology:面向数据智能体的自进化本体层)[02:48] 🤖 RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents(RecreationWorld:面向混合计算机使用智能体的可扩展且可验证环境)[03:28] 🧩 IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts(IntBMoE:将块级条件融入专家组合以实现全参与混合专家)[04:14] 🎥 OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue(OmniVChat:面向原生音视频对话的合成、基准测试与训练)[04:59] 🎨 Paint-Anything: Unified Any-Color Control for Image Generation and Editing(Paint-Anything:面向图像生成与编辑的统一任意颜色控制)[05:45] 🎬 OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation(OmniVBench:面向全能参考到视频生成的基准与大规模数据集)[06:25] 🧬 GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills(GraphSkillEvo:图结构智能体技能的进化优化)[07:06] 🤖 MintAct: A Unified Visual Agent for Digital Environments(MintAct:面向数字环境的统一视觉智能体)[07:49] 🎨 Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design(Designer-RSI:从用户流量中演化程序性记忆以支持智能体图形设计)[08:34] 🤖 When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation(当AI评审训练AI评审者:科学判断坍缩与缓解)[09:15] ⚖ Calibrating Teacher--Student Discrepancy for On-Policy Distillation(面向在线策略蒸馏的教师—学生差异校准)[09:57] 🛡 FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection(FRAUDSkill:面向音频反欺诈检测的结构化冻结权重技能优化)[10:41] 📞 TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection(TeleAntiFraud 2.0:一个可刷新、以画像为依据且基于音频的电信诈骗检测基准) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  5. Sep 18

    2026.09.18 | V4.1-Flash压缩KV提效;SoL-Pi优化智能体降本

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:32] 🗜 DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression(DeepSeek-V4.1-Flash:将 KV 缓存压缩推向极限)[01:15] ⚡ SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness(SoL-Pi:递归扩展自动化研究循环以实现高效智能体执行框架)[01:55] 🛑 When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation(当 EOS token 不一致:理解在线策略蒸馏中的长度膨胀)[02:33] 🧪 An Empirical Study of Harness Design for Coding Agents(面向编码智能体的 Harness 设计实证研究)[03:20] 🌍 JEPA-Anything: Learning Predictive Models across Different Worlds(JEPA-Anything:跨不同世界学习预测模型)[04:09] 🕵 RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation(RiskChainBench:面向混淆平台消息还原与证据支撑网络调查的基准)[04:56] 🎓 RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning(RetireOPD:面向智能体强化学习的自退场在线策略蒸馏)[05:40] 📄 WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing(WeVisDoc:从覆盖到能力,实现鲁棒的端到端文档解析)[06:29] 🔄 Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents(反思、修订、复用:面向GUI智能体的免训练技能演化)[07:10] 🤖 VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control(VABench:通过视觉演示、主动感知和度量控制测量具身空间智能)[07:56] 🎥 Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation(Video DeltaNet:面向直播视频生成的视频原生混合注意力)[08:42] 🦾 FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations(FAMOS:基于稀疏观测的前馈式三维铰接建模)[09:24] 🖼 UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation(UFO:面向多模态图像生成全条件对齐的评估链)[10:07] 🌍 Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model(MiniMax-H3 能否推理物理世界?一项全模态生成模型评估)[10:54] 🧠 When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models(When2Think:面向高效混合推理模型的难度感知长度控制学习) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  6. Sep 17

    2026.09.17 | 科学代码库转智能体环境;上下文机制网络赋能结构化数据智能

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:31] 🧪 ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments(ScienceIDE:将全球科学代码库转化为智能体可学习环境)[01:23] 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络)[02:10] 📉 Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening(重新思考PPO中的评论家学习:理解与缓解价值平坦化)[02:56] 🧠 Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents(置信度源于经验:从推理到智能体的经验性置信度估计)[03:36] 🤖 ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks(ProgramDistill:从交互式 Web 应用到可验证的参考引导软件工程任务)[04:21] 🤖 ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models(ActionPiece:重新思考自回归视觉-语言-动作模型的动作标记化)[05:01] 🧠 Agora: Git as Shared Memory for Collective AutoResearch(Agora:将 Git 作为集体自动研究的共享记忆)[05:43] ⚡ VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention(VC-Attention:面向低比特注意力的值平滑与 Softmax 转换)[06:22] 📈 EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents(EvolveTrade:面向自演化LLM交易智能体的经验驱动策略精炼)[07:02] 🎮 Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control(Zing-0.5:迈向具备实时联合动作与文本控制的可玩世界)[07:51] 👀 Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX(注视作为共同基础证据:MapTask与MUNDEX的跨语料库分析)[08:35] 🎯 A Zeroth-Order Paradigm for LLM Preference Alignment(面向LLM偏好对齐的零阶范式)[09:17] 🧬 HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses(HypoEvolve:遗传算法使多智能体大语言模型能够发现科学假设)[10:02] ✋ EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset(EventEgoHands++:基于事件的第一人称3D手部网格重建与真实数据集)[10:55] 🔭 SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization(SpectralShift:通过谱重参数化有效扩展 Gated DeltaNet 的上下文窗口) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  7. Sep 16

    2026.09.16 | 持续学习组合提升保留率;游戏AI六类角色待整合

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:29] 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆)[01:17] 🎮 AI for Games in the Foundation Model Era(基础模型时代的游戏人工智能)[02:05] 🎙 StepAudio 3 Realtime Technical Report(StepAudio 3 Realtime 技术报告)[02:45] 🎵 StepAudio 3 Music Technical Report(StepAudio 3 Music 技术报告)[03:22] 🤖 ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents(ScienceBuddy:面向交互式科学智能体的递归嵌套式自我改进)[04:09] 🧭 HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness(HarnessVLN:通过智能体框架统一免训练具身导航)[04:52] 🤖 The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement(人类建造的最后一个AI:迈向真正的递归自我改进)[05:40] 🏗 Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?(墙里的另一张蓝图:如何像孩子一样向前沿AI提问?)[06:26] 🧠 Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States(Mind2Dialogue:通过模拟用户心理状态训练人类感知语言模型)[07:10] 🧩 ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement(ModularRSI:模块化且可泛化的递归式执行框架自我改进)[07:54] 🤖 Modality-Autoregressive World-Action Models(模态自回归世界动作模型)[08:38] 📐 Disentangling Representation Evolution in Transformers through Directional Decomposition(通过方向分解解耦 Transformer 中的表征演化)[09:25] 🎥 PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control(PhysStream:具备结构化场景记忆与细粒度运动控制的流式物理基础视频生成)[10:06] 🧪 ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals(ImpossibleRubrics:对作为奖励信号生成的评分标准进行压力测试)[10:55] 🧠 Convergent Emergence of In-Context Learning Across Modalities(跨模态上下文学习的趋同涌现) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

Ratings & Reviews

5
out of 5
2 Ratings

About

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

You Might Also Like