HuggingFace 每日AI论文速递

duan

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

  1. 7h ago

    2026.08.19 | 进化策略微调省显存;技能应用有双刃剑

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:34] 🧬 Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements(Agentic ESOpt:以极低GPU需求微调长程LLM智能体)[01:46] 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效)[02:35] 🔬 ASI-Bench: At the Dawn of Artificial Superintelligence(ASI-Bench:人工超级智能的黎明)[03:26] 💻 FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution(FreeToken:高效的边缘原生MoE服务与带宽自适应执行)[04:20] 🧭 Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation(具身导航器:指向、思考、记忆与对齐实现高效导航)[05:22] 🎬 AVA-Encoder: Towards Agent-Native Video Representation Learning(AVA-编码器:迈向智能体原生的视频表示学习)[06:12] 🖼 EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing(EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑)[07:13] ⚡ Agent Lightning v1.0: Towards Harnessed Agentic RL(Agent Lightning v1.0:迈向框架化的智能体强化学习)[08:12] 🎥 V-RAE: Rethinking Video Latent Spaces for Generation(V-RAE:重新思考用于生成的视频潜空间)[09:17] 🎬 CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing(CoinVE-200K:面向组合式指令引导视频编辑的大规模高质量数据集)[10:17] 🛡 DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization(DiSCO:通过分布引导的对比提示优化防御文本到图像生成)[11:22] 🧠 Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents(驾驭记忆:对记忆智能体中记忆底层介质的整体评估)[12:32] 🎨 From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation(从语料到协同演进的能力:以能力为中心的通用图像生成数据设计)[13:31] 📊 StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows(StartupBench:对通用智能体在市场验证的端到端工作流上的基准测评)[14:34] ⚡ Energy-Guided Flow Matching(能量引导的流匹配) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  2. 1d ago

    2026.08.18 | 智能体化评测让世界模型可诊断;多模态三维生成仍有瓶颈

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:32] 🕵 HarnessEval-W: Agentifying the Evaluation of Visual Worlds(HarnessEval-W:使视觉世界的评估智能体化)[01:32] 🌍 VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?(VibeWorlding:多模态智能体能否端到端构建3D开放世界?)[02:10] ⚡ MOSS-VL Technical Report(MOSS-VL 技术报告)[03:05] 🤖 ClawGym II: Exploring Black-Box RL on Agent Harness(ClawGym II:在智能体框架上探索黑盒强化学习)[03:56] 🎯 Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization(学习尚未掌握的,而非已经精通的:面向多奖励策略优化的饱和感知优势重加权)[04:51] 🤖 UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations(UI-Mate:利用上下文演示推进开放权重的基础图形用户界面智能体)[05:46] 🔬 Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search(大型发现模型:基于经验建模的开放式搜索)[06:37] 🎨 An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models(训练像素空间文本到图像扩散模型的实证研究)[07:31] 🤖 Agentic Transaction: Towards ACID-Compliant Agent Systems(智能体事务:迈向ACID合规的智能体系统)[08:29] 🔬 How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks(智能体如何在自动研究中失败:基于100个真实前沿研究任务的端到端诊断评估)[09:24] ⚡ GenRouter: Unified Workflow Routing for Agentic Image Generation(GenRouter:用于智能体图像生成的统一工作流路由)[10:29] 🧩 MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling(MegaParts:通过词元高效的自回归建模将部件感知的3D物体生成扩展到300个部件)[11:22] 🧠 Understanding Cognition-Induced Risks in Agentic AI Systems(理解智能体AI系统中认知引发的风险)[12:21] 🔗 Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI(推进开放可复现的关系学习:RelArena-α、TabPFN-Rel与RPI)[13:24] 🛡 Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs(Ventor-QTest:威胁模型驱动的供应商托管LLM API验证) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  3. 2d ago

    2026.08.17 | 视频检测难防伪;自监督蒸馏促提升

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:30] 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估)[01:27] 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏)[02:21] 🤖 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development(超越最终得分:对长周期AI研究与开发智能体的系统评估)[03:15] 🧠 Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning(Intern-S2-Mobius:知识与推理解耦的基础模型)[04:09] 🎮 Marionette: Predicting World States, Rendering Geometry, Painting Appearance(Marionette:预测世界状态,渲染几何,绘制外观)[05:00] 🧠 SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning(SimpleOPD:面向长上下文推理的简单分词器无关在线策略蒸馏)[06:03] 🧠 MobileMem: Learning from a Year of Mobile Experiences(移动记忆:从一年的移动体验中学习)[06:54] 🤖 DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data(DFM Mimir v1:仅使用合规后训练数据、以1B参数实现前沿性能的开放HRM)[07:50] 🤸 HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark(HumanTracker:迈向全面且与人类感知一致的运动追踪基准)[08:50] 🎨 CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing(CPI-Bench:一个面向真实世界图像编辑的全面、实用且智能的基准)[09:48] 🧠 Latent On-Policy Self-Distillation(潜在同策略自蒸馏)[10:53] 🔍 Claim-Level Reliability Assessment for Efficient Test-Time Reasoning(面向高效测试时推理的声明级可靠性评估)[11:43] 🤖 PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment(PRM-as-a-Judge 1.5:机器人过程评估工具包)[12:42] 📉 Forecast Collapse in Time-Series Foundation Models(时间序列基础模型中的预测崩溃)[13:41] 🤔 Second Thought: Reasoning in Parallel as LLM Agents Act and Observe(第二思考:LLM智能体在行动与观察时并行推理) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  4. 5d ago

    2026.08.14 | 动作条件视频世界模型引入几何感知;长时记忆外部化实现无尽世界

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:32] 🤖 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation(DreamX-Phi 1.0:面向机器人操作的动作条件视频世界模型)[01:24] 🌍 Alaya-EVOKE: From Linear-Scaling Supervision to Endless World(Alaya-EVOKE:从线性扩展监督到无尽世界)[02:30] 🔀 LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers(LLMRouter:开发、评估和部署LLM路由器的统一基础设施)[03:25] 🧬 DarwinX: Evolving Agent Harnesses Through Natural Selection(DarwinX:通过自然选择进化智能体框架)[04:26] 🔬 Intern-S2-Preview: Scientific Agentic Foundation Model(Intern-S2-Preview:科学智能体基础模型)[05:20] 🎮 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives(PlayWorld:使用智能体玩家在长程目标上对世界模型进行基准测试)[06:20] 🤖 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design(AutoDesign:面向长时程智能体设计的元框架优化)[07:23] 🧠 Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence(空间记忆智能体:基于经验的程序记忆实现空间智能)[08:12] ⚡ Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus(混合线性注意力大语言模型中的大规模激活:注意力前尖峰与尖峰间平台)[09:03] 🎭 UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos(UniSwap:面向说话视频的流式音频-视觉身份交换)[10:12] ⚡ LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time(LiveAnimate:实时稳定长格式流式人体动画生成)[11:08] 🤖 How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review(修辞何以能对AI审稿人进行奖励黑客?解析基于AI的同行评审中的修辞敏感性)[12:00] ✂ An AI4AI Framework for Visual Token Pruning(面向视觉Token剪枝的AI4AI框架)[12:56] 🤖 H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models(H2R-Bench:在世界模型中评估人类到机器人的操作视频生成)[14:08] 🔄 Full-bandwidth transformer(全带宽Transformer) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  5. 6d ago

    2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:26] 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试)[01:35] 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成)[02:43] 🧩 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses(测试时AI4AI:通过推理支架实现强到弱能力迁移)[03:37] 🔬 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence(Mechanist:将人工智能作为揭示智能机制的科学仪器)[04:33] 🎭 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives(大语言模型智能体能否坚守剧本?交互式叙事中长程一致性的基准测试)[05:33] 🌍 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization(StateFlow:为预可视化构建、演化与访问3D世界状态)[06:28] 📐 Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models(自几何:面向几何一致的3D视觉基础模型的无真值即插即用测试时自适应)[07:25] 🪞 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection(从合成到去除:物理驱动的反射模拟与基于扩散模型的视频去反射)[08:22] 🛡 ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents(ToolHazard:扩展对抗性环境,用于基于大语言模型的智能体的安全评估与对齐)[09:15] 🔍 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images(视觉工具使用的幻象:图像思维的因果审计)[10:16] 🤖 Self-Evolving Embodied Agents via Skill-Harness Evolution(基于技能与执行框架演化的自进化具身智能体)[11:19] 🧩 SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries(SkillZip:面向可扩展智能体技能库的合约保持图压缩)[12:14] 🛡 Agent Safety Should Be a Runtime Contract(智能体安全应当是一种运行时契约)[13:01] 🧬 Persistent Recursive Worlds Enable Autonomous Software Evolution(持久递归世界赋能自主软件演化)[13:46] 💡 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation(MBA:面向真实世界商业构思的多模态基准与智能体) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  6. Aug 12

    2026.08.12 | 共体智能体以人为中心助人成长;智能体与环境共演化迈向自我导向

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:35] 🤖 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI(共体智能体:以人为中心的智能体人工智能新范式)[01:22] 🧬 Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design(智能体系统中的共同演化:迈向超越人类设计的自我导向演化)[02:22] 🌍 Beyond Pixels: From Video Priors to 4D Worlds(超越像素:从视频先验到4D世界)[03:12] 🧩 Articulated Object Reconstruction from Rest-State Observation(基于静止状态观测的铰接物体重建)[04:09] ⚔ AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss(AdvFD:通过对抗性弗雷歇距离损失提升视觉生成)[05:05] 🧬 Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution(孟德尔·哥德尔机:通过比较进化实现递归自我改进的编码智能体)[06:00] 🎭 Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence(Ex-Omni-2D:具备原生视觉临场感的表现性全模态对话模型)[06:48] 🌍 VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?(VibeLifeBench:你的生活智能体能否在动态世界中主动且持久地行动?)[07:46] 🚫 Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness(解码级禁忌:大语言模型鲁棒性的诊断性压力测试)[08:33] 📦 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure(SkillZip:通过发现可复用结构实现自我进化智能体的免评估技能压缩)[09:33] 📱 SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information(SPIEval:评估大型语言模型作为移动助手处理分散个人信息的能力)[10:36] ✂ Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents(并不值得再投入一个词元:高效深度研究智能体的边际价值估计)[11:34] 🌐 Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation(开放大语言模型用于多语言机器翻译的无参考后训练)[12:28] 🔍 InSight-doc: Agentic Visual Perception for Long-Document Understanding(InSight-doc:面向长文档理解的智能体视觉感知)[13:27] 🔀 UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models(UniMoMo:基于专家合并的大型推荐模型MoE加速方法) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  7. Aug 11

    2026.08.11 | 自进化混合专家赋能持续学习;代码重构基准揭示智能体局限

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:30] 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习)[01:24] 🔧 SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring(SWE-Bench ProMax:面向大规模多语言代码重构的智能体基准评测)[02:21] 🐍 Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution(Ouroboros:通过核心评审进化实现自我发展的前沿编程智能体)[03:31] 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习)[04:13] 🧠 Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory(智能体记忆蒸馏:利用分层教师记忆赋能小型大语言模型智能体)[05:11] 🧠 Motif 3: Technical Report(Motif 3:技术报告)[05:59] 🔬 Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains(Sci-VBench:评估科学领域中知识与推理密集型视频生成)[06:49] 🖼 What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems(下一步编辑什么:对话系统中的视觉对齐图像编辑后续建议)[07:53] 🎯 SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation(SPOT:面向同策略蒸馏的稀疏探测与结果校准)[08:58] ⚡ OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching(OasisKV:通过前瞻稀疏预取将解码期KV缓存扩展到HBM之外)[09:49] 🧠 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States(RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱)[10:43] 🔍 Evidence-RL: Towards Evidence-intensive Visual Reasoning(证据强化学习:迈向证据密集型视觉推理)[11:42] 🧠 Scaling Inherently Interpretable Language Models(扩展内在可解释的语言模型)[12:40] 🧬 Evo-Bench: Can Language Models Improve Agent Harness?(Evo-Bench:语言模型能否改进智能体运行框架?)[13:40] 🔓 Stealing Reasoning Traces from Proprietary LLM APIs(从专有大语言模型API中窃取推理轨迹) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

Ratings & Reviews

5
out of 5
2 Ratings

About

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

You Might Also Like