Search papers, labs, and topics across Lattice.
This paper introduces CoinVE-200K, a large-scale dataset designed for compositional instruction-guided video editing, addressing the limitations of existing datasets that focus on single editing operations. The dataset includes 1080p video-editing pairs with 2 to 5 atomic editing operations, ensuring high-quality, temporally consistent, and visually diverse outputs. Additionally, the authors present CoinVE-Edit, a 22B model that utilizes region-aware attention to enhance multi-region editing capabilities, achieving superior performance on the newly established CoinVE-Bench benchmark.
CoinVE-200K enables video editing models to understand and execute complex, multi-faceted editing instructions with unprecedented accuracy and quality.
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.