Search papers, labs, and topics across Lattice.
This paper introduces AVE-Compass, a novel benchmark designed to evaluate audio-video editing capabilities by addressing the intertwined nature of audio and visual signals in real-world videos. The benchmark comprises 145 curated videos and 196 editing instructions, assessing models on Instruction Following, Fidelity Preserving, Realism, and Editing Intent through a combination of checklist-based evaluations and automated metrics. Results reveal that current state-of-the-art models struggle with cross-modal instructions, prompting the development of AVE-Agent, a modular framework that enhances editing performance by breaking down complex tasks and leveraging self-reflection and evaluator feedback.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.