HuggingFace 每日AI论文速递

duan

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

  1. 23h ago

    2026.08.24 | 大模型超参迁移降本;图工程引领系统协作

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:32] 🚀 Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts(让我们一步一步扩展规模:面向大规模混合专家模型的高效超参数迁移)[01:24] 🕸 Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence(大语言模型智能体时代的图工程:从个体智能到系统智能)[02:21] 🤖 OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs(OmniAssistBench:全模态大语言模型的助手式交互基准)[03:17] ⚡ ParaTempo: Efficient Parallel Reasoning via Temporal Confidence(ParaTempo:基于时间置信度的高效并行推理)[04:00] ♾ InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter(InfinityEdit:基于轻量级编辑触发适配器的无限视频编辑)[04:53] 🪙 Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models(每一枚硬币都有两面:论大型语言模型同策略蒸馏中泛化的双重性)[05:37] 🧩 EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking(EviRank:面向多模态图像重排序的结构化相关性证据)[06:46] 🧠 Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs(超越正确性:混合思考多模态大语言模型的响应行为基准测试与对齐)[07:45] 🎨 UniSpace: Unified Visual Representation and Scalable Multimodal Modeling(UniSpace:统一视觉表示与可扩展多模态建模)[08:39] 🛒 Towards Faithful Simulation of Human Shopping Behavior(面向人类购物行为的忠实模拟)[09:35] ⚙ AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale(AgentMercury:你的智能体能够大规模合成可验证的商业场景环境)[10:41] ⚡ Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference(Daedalus-150M:为CPU推理设计的卷积-注意力混合模型)[11:35] 🛡 CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment(CLEAR:面向保持实用性的大语言模型安全对齐的连续潜在适配器路由)[12:42] 📱 Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs(Llama-Mobile:高效2.7比特视觉语言模型量化)[13:32] 🍳 FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth(FlavourBench:用可执行的烹饪真值对前沿语言模型进行排名) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  2. 3d ago

    2026.08.21 | 动态环境适配弱点;任务合成保留源意图

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:31] 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界)[01:28] 🖥 FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis(FACET:在终端任务合成中保留源意图与可执行状态)[02:29] 🧪 SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?(SWE-bench Science:编码智能体能否解决科学领域的工程任务?)[03:15] 🕺 4DAnyone: Create Anyone in 4D from a Casual Monocular Video(4DAnyone:从随意单目视频创建任意人物的4D形象)[04:11] 👥 WithEveryone: Unified Planning and Identity Grounding for Group Image Generation(WithEveryone:面向群像生成的统一规划与身份锚定)[05:01] 🧠 MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use(MemTrapBench:大语言模型记忆使用中的认知陷阱基准测试)[05:52] 🔄 SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback(SkillEvo:来自多轮交互反馈的自我更新进化梯度)[06:50] 🎮 ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models(ForgeWM:面向少步动作条件视频世界模型的渐进式因果训练)[07:48] 🧩 Repo0: Design-Driven Zero-to-All Code Generation(Repo0:设计驱动的从零到全代码生成)[08:46] ⚡ FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving(FlashPrefill V2:面向长上下文大语言模型服务的块稀疏预填充注意力)[09:41] 🧠 Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization(注入、对齐、恢复:面向无检索文档知识内化的分阶段后训练)[10:42] 🤖 EXIMO: VLM Guided Exploration of VLA Policies(EXIMO:视觉语言模型引导的VLA策略探索)[11:36] 🧠 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See(用低资源语言思考:SFT构建了什么,RL修复了什么,准确率无法看到什么)[12:18] 🎯 Towards Quantifying Benchmark Optimization in ASR Models(面向ASR模型中基准优化的量化研究)[13:20] 🛡 PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents(PolicyGuide:从守护单一动作到引导策略合规型LLM智能体的整个工作流) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  3. 4d ago

    2026.08.20 | 闭环进化提升具身智能;验证门控保障工业代码

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:34] 🤖 Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence(Zetta ζ:面向自进化物理智能的高效闭环具身智能体框架)[01:29] ✅ SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation(SemaPLC:一种基于项目、以验证为门控的PLC代码生成智能体框架)[02:31] 🎯 SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation(SemComp-Bench:视频生成中语义任务完成的基准测试)[03:28] 🔬 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist(全能科学家:一个全模态、全学科的人工智能科学家)[04:26] 🧠 Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL(Co-RL:多智能体强化学习中多样化群体催生无监督推理)[05:13] 🎮 SPADE: Self-Play in Adaptive Synthetic Executable Environments(SPADE:自适应合成可执行环境中的自博弈)[06:08] 🧪 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis(训练面向单步逆合成的化学合理性感知大语言模型)[07:04] 🎯 Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning(潜在世界模型中的决策度量对齐:诊断与面向MPC规划的动作条件目标)[07:58] 🧬 Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification(训练留下痕迹:面向语言模型谱系验证的中心化残差签名)[08:50] 🔁 Looped Language Models Improve Compositional Tool Calling(循环语言模型提升组合式工具调用能力)[09:36] 🖐 SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation(SoftVTBench:面向可变形物体操作的变形感知视触觉数据集与基准)[10:33] ⚽ FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents(FM-Bench:面向竞争智能体的长时程管理基准)[11:25] ✍ Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion(借助属性引导的体裁扩展,将创意写作扩展到故事中心数据之外)[12:15] 🔥 The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning(越热门越难遗忘:大语言模型遗忘的自适应流行度方法)[13:01] 🔍 Temporal Multi-Signal Fusion for Token-Level Hallucination Detection(面向Token级幻觉检测的时序多信号融合) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  4. 5d ago

    2026.08.19 | 进化策略微调省显存;技能应用有双刃剑

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:34] 🧬 Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements(Agentic ESOpt:以极低GPU需求微调长程LLM智能体)[01:46] 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效)[02:35] 🔬 ASI-Bench: At the Dawn of Artificial Superintelligence(ASI-Bench:人工超级智能的黎明)[03:26] 💻 FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution(FreeToken:高效的边缘原生MoE服务与带宽自适应执行)[04:20] 🧭 Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation(具身导航器:指向、思考、记忆与对齐实现高效导航)[05:22] 🎬 AVA-Encoder: Towards Agent-Native Video Representation Learning(AVA-编码器:迈向智能体原生的视频表示学习)[06:12] 🖼 EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing(EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑)[07:13] ⚡ Agent Lightning v1.0: Towards Harnessed Agentic RL(Agent Lightning v1.0:迈向框架化的智能体强化学习)[08:12] 🎥 V-RAE: Rethinking Video Latent Spaces for Generation(V-RAE:重新思考用于生成的视频潜空间)[09:17] 🎬 CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing(CoinVE-200K:面向组合式指令引导视频编辑的大规模高质量数据集)[10:17] 🛡 DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization(DiSCO:通过分布引导的对比提示优化防御文本到图像生成)[11:22] 🧠 Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents(驾驭记忆:对记忆智能体中记忆底层介质的整体评估)[12:32] 🎨 From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation(从语料到协同演进的能力:以能力为中心的通用图像生成数据设计)[13:31] 📊 StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows(StartupBench:对通用智能体在市场验证的端到端工作流上的基准测评)[14:34] ⚡ Energy-Guided Flow Matching(能量引导的流匹配) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  5. 6d ago

    2026.08.18 | 智能体化评测让世界模型可诊断;多模态三维生成仍有瓶颈

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:32] 🕵 HarnessEval-W: Agentifying the Evaluation of Visual Worlds(HarnessEval-W:使视觉世界的评估智能体化)[01:32] 🌍 VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?(VibeWorlding:多模态智能体能否端到端构建3D开放世界?)[02:10] ⚡ MOSS-VL Technical Report(MOSS-VL 技术报告)[03:05] 🤖 ClawGym II: Exploring Black-Box RL on Agent Harness(ClawGym II:在智能体框架上探索黑盒强化学习)[03:56] 🎯 Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization(学习尚未掌握的,而非已经精通的:面向多奖励策略优化的饱和感知优势重加权)[04:51] 🤖 UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations(UI-Mate:利用上下文演示推进开放权重的基础图形用户界面智能体)[05:46] 🔬 Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search(大型发现模型:基于经验建模的开放式搜索)[06:37] 🎨 An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models(训练像素空间文本到图像扩散模型的实证研究)[07:31] 🤖 Agentic Transaction: Towards ACID-Compliant Agent Systems(智能体事务:迈向ACID合规的智能体系统)[08:29] 🔬 How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks(智能体如何在自动研究中失败:基于100个真实前沿研究任务的端到端诊断评估)[09:24] ⚡ GenRouter: Unified Workflow Routing for Agentic Image Generation(GenRouter:用于智能体图像生成的统一工作流路由)[10:29] 🧩 MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling(MegaParts:通过词元高效的自回归建模将部件感知的3D物体生成扩展到300个部件)[11:22] 🧠 Understanding Cognition-Induced Risks in Agentic AI Systems(理解智能体AI系统中认知引发的风险)[12:21] 🔗 Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI(推进开放可复现的关系学习:RelArena-α、TabPFN-Rel与RPI)[13:24] 🛡 Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs(Ventor-QTest:威胁模型驱动的供应商托管LLM API验证) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

  6. Aug 17

    2026.08.17 | 视频检测难防伪;自监督蒸馏促提升

    【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】本期的 15 篇论文如下: [00:30] 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估)[01:27] 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏)[02:21] 🤖 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development(超越最终得分:对长周期AI研究与开发智能体的系统评估)[03:15] 🧠 Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning(Intern-S2-Mobius:知识与推理解耦的基础模型)[04:09] 🎮 Marionette: Predicting World States, Rendering Geometry, Painting Appearance(Marionette:预测世界状态,渲染几何,绘制外观)[05:00] 🧠 SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning(SimpleOPD:面向长上下文推理的简单分词器无关在线策略蒸馏)[06:03] 🧠 MobileMem: Learning from a Year of Mobile Experiences(移动记忆:从一年的移动体验中学习)[06:54] 🤖 DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data(DFM Mimir v1:仅使用合规后训练数据、以1B参数实现前沿性能的开放HRM)[07:50] 🤸 HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark(HumanTracker:迈向全面且与人类感知一致的运动追踪基准)[08:50] 🎨 CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing(CPI-Bench:一个面向真实世界图像编辑的全面、实用且智能的基准)[09:48] 🧠 Latent On-Policy Self-Distillation(潜在同策略自蒸馏)[10:53] 🔍 Claim-Level Reliability Assessment for Efficient Test-Time Reasoning(面向高效测试时推理的声明级可靠性评估)[11:43] 🤖 PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment(PRM-as-a-Judge 1.5:机器人过程评估工具包)[12:42] 📉 Forecast Collapse in Time-Series Foundation Models(时间序列基础模型中的预测崩溃)[13:41] 🤔 Second Thought: Reasoning in Parallel as LLM Agents Act and Observe(第二思考:LLM智能体在行动与观察时并行推理) 【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递 在小宇宙查看该单集文稿

Ratings & Reviews

5
out of 5
2 Ratings

About

每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】

You Might Also Like