HuggingFace 每日AI论文速递

duan

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

  1. 1d ago

    2026.10.02 | 流式视频主动记忆;音视频扩散奖励路由

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:29] 🎥 OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction(OneStreamer:统一流式视频交互中的感知、记忆与主动响应)[01:11] 🔀 Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL(自适应奖励路由:通过前向过程强化学习实现联合音视频扩散的动态多奖励优化)[01:55] 🧠 Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States(超越记忆:利用显式信念状态驾驭长时程智能体)[02:38] 🤖 Agent Priors-guided Policy Learning(智能体先验引导的策略学习)[03:25] 🌀 Hierarchical Continuous Diffusion Language Models(分层连续扩散语言模型)[04:03] 👁 World Observer: Joint Actor-Observer Generation for Persistent World Modeling(World Observer:面向持久世界建模的联合行动者-观察者生成)[04:54] 📉 Sharpening Tax in Post-Training(后训练中的锐化税)[05:39] 🎯 ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization(ActiveSaddler:面向智能体执行框架优化的自动化课程学习)[06:22] 🤖 A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review(可信AI审稿人缺失的一环:从修辞鲁棒性基准测试到SciCore评审)[07:08] 🎮 ROWBench: Do Video Models Render What the Program Specifies?(ROWBench:视频模型能否渲染程序所指定的内容?)[07:55] 🤖 AutoGUIWorld: Image Generators as Visual World Models for GUI Agent(AutoGUIWorld:将图像生成器用作 GUI 智能体的视觉世界模型)[08:39] 🔍 Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation(基于跨执行框架适配的检索增强技能优化)[09:23] ⚖ Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL(让稀疏奖励算数:多奖励强化学习中的密度感知奖励聚合)[10:05] 🤖 Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding(去中心化 Master-Mind:多智能体路径规划中通过迭代意图去噪的联合动作精炼)[11:02] 🧩 E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models(E-MoE:面向非因子化扩散语言模型的增强混合专家) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  2. 2d ago

    2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:27] 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差)[01:11] 🪞 UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement(UniEvo-VL:面向多模态模型自我改进的在线策略自蒸馏训练方案)[01:54] 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊)[02:42] 🤖 AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks(AREX-2:通过长时程反思任务推进自我改进智能体)[03:29] 🖥 Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents(Mid-Harness:在模型与执行框架之间扩展终端智能体的动作)[04:09] 🧬 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery(EvoDuet:面向科学发现的网络搜索与任务求解双层协同演化)[05:01] 🕵 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents(WorldAuditBench:使用多模态智能体进行交互式3D世界审计)[05:47] 🛠 Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI(测试时 AI4AI 中面向智能体执行框架设计的元技能学习)[06:26] 🧠 EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making(EVOKE:激发智能体中的世界知识以实现可迁移决策)[07:12] 🎮 RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement(RSIGame:具备递归自我改进能力的自主智能体游戏开发)[07:53] 📉 More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models(更多选择,更少决策:类JEV直接决策模型中的序数尺度偏差)[08:45] 🧠 Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering(Imagine3D-LLM:教会多模态大语言模型在回答前想象3D场景)[09:31] 🛠 Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training(智能体错误数据集:面向失败分析与错误感知后训练,规模化构建5万条错误—诊断配对)[10:21] 🏮 LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models(LANTERN:照亮语言模型中的隐藏数学知识)[11:05] 🖼 It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them(并非图像所示:无关上下文会扰乱 VLM 评判模型却不为其提供信息) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  3. 2d ago

    【月末特辑】9月最火AI论文 | LimiX-2结构化智能;Vidu S2实时可编辑空间视频

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 10 篇论文如下: [00:38] TOP1(🔥810) | 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络)[02:36] TOP2(🔥704) | 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成)[04:41] TOP3(🔥494) | 🤖 Raven: The Harness of Harnesses for Composable Agentic Intelligence(Raven:面向可组合智能体智能的“框架之框架”)[06:43] TOP4(🔥493) | 🎓 StudentSim: Training LLM-based Student Simulators(StudentSim:训练基于大语言模型的学生模拟器)[08:51] TOP5(🔥484) | 🤖 Scaling Automatic Research Agents via World Models(通过世界模型扩展自动研究智能体)[11:10] TOP6(🔥431) | 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明)[13:35] TOP7(🔥396) | 🤖 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills(从仓库到技能:将GitHub代码库蒸馏为AI4AI技能)[15:38] TOP8(🔥382) | 🎨 MaLiang-Harness: A Programmable Path to Image and Video Generation(MaLiang-Harness:通往图像与视频生成的可编程路径)[17:42] TOP9(🔥376) | 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆)[19:55] TOP10(🔥358) | 🤖 In-Context Learning for Robots: Methods and Applications(面向机器人的上下文学习:方法与应用) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  4. 5d ago

    2026.09.28 | 层融合正则缩小重建生成差;血缘数据流加速保序容错

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:28] 🧩 FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders(FuseReg:正则化层融合缓解表征自编码器中的重建-生成差距)[01:12] 🔗 RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation(RayOrch:面向基础模型数据准备的受血缘控制多粒度数据流编程与执行)[01:58] ⚡ Block Sparse Attention with Log-Linear Complexity(对数线性复杂度的块稀疏注意力)[02:40] 🤖 InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data(InternW0-Δ:一个连接预测动态与动作、基于20K+小时开放数据的世界动作模型)[03:24] 🖐 Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors(Tactile-JEPA:面向分布式触觉传感器的拓扑感知自监督表示学习)[04:11] 🛰 Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning(利用预训练扩散模型与多模态条件增强摄影测量数字表面模型)[04:58] 👁 FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance(FoMo:生成轨迹中的分叉时刻作为感知距离)[05:42] 📊 Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem(真实环境中的 Jev:Jev 模型功能、应用与生态的数据驱动分析)[06:30] 🎯 TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations(TrackEverything:通过去重3D场景表示实现长时程密集跟踪)[07:21] 🧩 SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL(SLCA-GRPO:解决工具调用强化学习中的跨段信用误归因)[08:08] 🎯 CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation(CARD:面向个性化文本生成的聚类级适配与奖励引导解码)[08:54] ⚖ Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs(隐式个性化与显式风格会冲突吗?PsPLUG:用于平衡定制化 LLM 中个性化与风格的轻量级插件)[09:41] 🤝 AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs(AgentWorld:多智能体大语言模型长时程协作基准测试)[10:24] 🎮 Game Arena: Strategic LLM Evaluation in Competitive Environments(游戏竞技场:竞争环境中的大语言模型策略评估)[11:08] 🏦 IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking(IndicBankBench:评估印度零售银行中语言模型助手的安全性与可靠性) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  5. Sep 26

    2026.09.25 | 世界模型物理推理可评测;大模型线性叠加双想法

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:27] 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性)[01:12] 🧠 Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs(你的Transformer能同时容纳两个想法:LLM中线性叠加的证据)[01:53] 🎬 WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation(WanPE:面向现代文本到视频生成的电影级提示增强)[02:40] 🧭 OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents(OmniEcho:面向全模态具身智能体的音视频空间理解)[03:18] 🤖 Agent-Editing World Model: Rethinking World Modeling for LLM Agents(智能体编辑世界模型:为LLM智能体重新思考世界建模)[04:05] 🧩 Parts-of-Speech as Emergent Categories in SAE Latent Space(词性作为SAE潜空间中的涌现类别)[04:46] 🤖 Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents(Qwen-Planner-Agent:面向真实世界移动规划智能体的闭环 AI-for-AI 框架)[05:20] 🔍 IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis(IterSynth:通过角色解耦的迭代合成重新思考深度搜索智能体)[06:04] 🤖 Coding Agents for Generalized Task and Motion Planning Problems(面向泛化任务与运动规划问题的编码智能体)[06:49] 🧠 Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone(神经谱容量:仅凭网络规格测量与设计架构)[07:28] 🛡 AgentKernel: The Trust-Native Agentic Operating System(AgentKernel:信任原生的智能体操作系统)[08:15] 🧪 Rufus-Air: An Open LLM Post-Training Recipe(Rufus-Air:开放的大语言模型后训练配方)[08:59] 🤖 PUBG Ally: A Conversational Embodied Agent as an AI Teammate(PUBG Ally:作为AI队友的对话式具身智能体)[09:44] 🤖 World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal(世界动作智能体:利用视觉语言模型通过世界动作预演实现机器人操作)[10:27] 🛸 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds(ExplorationBench:测量 AI 系统在可验证异星世界中的探索能力) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  6. Sep 25

    2026.09.24 | 说话者双轨记忆;交互学习空间推理

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:28] 🧠 SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue(SpeakerMem-R1:以说话者为中心的多方对话双轨记忆)[01:13] 🤖 Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World(Spatial-Interactor:通过与可观测物理世界交互学习空间推理)[01:59] 🌍 HappyWorld-Bench(快乐世界基准(HappyWorld-Bench))[02:40] 🧠 The Past Frames the Future: Memory for Autoregressive Video Generation(过往帧塑造未来:自回归视频生成中的记忆机制)[03:29] 🧠 Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents(即时记忆:学习为LLM智能体整理任务自适应记忆)[04:11] 🎬 RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling(RewardVerse:面向视频奖励建模的评分量规引导策略优化)[05:00] 🎯 PACT: From Credit Assignment to Critic Alignment(PACT:从信用分配到评论家对齐)[05:43] 🧪 Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?(薛定谔的代码仓库:LLM 是学会了 SWE-bench,还是记住了它?)[06:24] 📐 GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression(GeoPair:用于免训练 Transformer 压缩的几何保持跨层因子分解)[07:00] 📦 PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing(PackLab:在机器人装箱中开发、训练与评估多模态大语言模型的综合框架)[07:53] 🧠 MemBodied: Recurrent Associative Memory for Vision-Language-Action Models(MemBodied:面向视觉-语言-动作模型的循环联想记忆)[08:39] 🧪 WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents(WhatWorkedBench:基准测试AI智能体的实验理解能力)[09:21] 🧠 Hunyuan-A13B Technical Report(混元-A13B 技术报告)[10:03] 🎮 Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms(可验证隐藏动力学游戏:从已求解机制生成智能体强化学习环境)[10:45] 🎥 All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation(所有模态都是平等的,但视频更平等:弥合联合视频生成中的交叉注意力差距) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  7. Sep 22

    2026.09.22 | VLM智能迁移至机器人控制;隐式3D记忆构建视频世界模型

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:28] 🤖 Transferring the Intelligence of VLMs to Robotic Control(将VLM的智能迁移至机器人控制)[01:11] 🎥 WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory(WorldCrafter:具有隐式3D感知记忆的一致视频世界模型)[01:53] 🎮 GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay(GameHorizon 套件:游戏玩法中的多时间跨度数据与评估)[02:35] 🧬 RRSI: Regularized Recursive Self-Improvement of Agent Harnesses(RRSI:智能体运行框架的正则化递归自我改进)[03:20] 📄 Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion(文档检索感知分块(D-RAC):通过 PDF 规范化与多模态 Markdown 转换实现企业文档的通用检索感知摄取)[04:01] 🎓 OmniEdu: Open Foundation Models for Learning and Teaching(OmniEdu:面向学习与教学的开放基础模型)[04:47] 🎬 VideoGen-Agent: Reinforcing Video Generation Agents(VideoGen-Agent:强化视频生成智能体)[05:37] 🐼 onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction(onPanda:通过Token级纠正为LLM与智能体高效标注同策略对齐数据)[06:19] 🤖 Grounded Action Model: 3D Grounding as a Foundation for Robotics(接地动作模型:以3D接地作为机器人学基础)[07:05] 🤖 One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents(从一到多,从多到一:面向软件工程智能体的类别感知迭代专家训练)[07:57] 🎭 Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations(Deep Persona:面向角色扮演智能体与模拟的心理学基础架构与评估框架)[08:36] 🧠 Harness-Zero: Harness Distillation via Agent-as-Harness(Harness-Zero:通过智能体作为外壳进行外壳蒸馏)[09:18] 🤖 CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies(CARE:面向视觉-语言-动作策略的经验引导式原子纠正执行)[10:13] 🧠 Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents(Jev-Mem:面向高效 AI 智能体的 System-One 控制型智能体记忆)[10:52] 🎥 Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms(为什么视频扩散模型会违背物理规律?揭示注意力机制中的缺陷) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

Ratings & Reviews

5
out of 5
2 Ratings

About

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

You Might Also Like