10月27日
单集 1750
7 分钟

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

In this episode, we discuss Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset by Qingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu, Hanlin Wang, Wen Wang, Ka Leong Cheng, Shuailei Ma, Yanhong Zeng, Zichen Liu, Yinghao Xu, Yujun Shen, Qifeng Chen. The paper presents Ditto, a comprehensive framework that generates large-scale, high-quality training data for instruction-based video editing by combining an advanced image editor with an in-context video generator. Ditto uses an efficient, distilled model with a temporal enhancer and an intelligent agent to ensure scalable, diverse, and high-fidelity video edits. Leveraging this framework, the authors created the Ditto-1M dataset and trained the Editto model, achieving state-of-the-art performance in following editing instructions.

单集网页

节目

AI Breakdown
频率

一日一更
发布时间

2025年10月27日 UTC 04:44
长度

7 分钟
单集

1750
分级

儿童适宜

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

信息