5 DAGE SIDEN
EPISODE 1,8 T
7 MIN.

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

In this episode, we discuss Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset by Qingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu, Hanlin Wang, Wen Wang, Ka Leong Cheng, Shuailei Ma, Yanhong Zeng, Zichen Liu, Yinghao Xu, Yujun Shen, Qifeng Chen. The paper presents Ditto, a comprehensive framework that generates large-scale, high-quality training data for instruction-based video editing by combining an advanced image editor with an in-context video generator. Ditto uses an efficient, distilled model with a temporal enhancer and an intelligent agent to ensure scalable, diverse, and high-fidelity video edits. Leveraging this framework, the authors created the Ditto-1M dataset and trained the Editto model, achieving state-of-the-art performance in following editing instructions.

Episodewebside

Serie

AI Breakdown
Hyppighed

Dagligt
Publiceret

27. oktober 2025 kl. 04.44 UTC
Længde

7 min.
Episode

1,8 t
Vurdering

Ikke anstødeligt

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

Oplysninger