Search papers, labs, and topics across Lattice.
This paper introduces OmniVideo-100K, a novel dataset designed for audio-visual Question Answering (QA) that addresses the limitations of existing video-caption-QA paradigms. By employing an automated data engine that includes Entity-Anchored Video Scripting and Clue-Guided QA Generation, the authors ensure coherent cross-segment referential consistency and enhance the depth of cross-modal reasoning. Fine-tuning state-of-the-art models on this dataset resulted in performance improvements of up to 20.59% on the human-verified test set, demonstrating significant advancements in audio-visual reasoning capabilities.
Fine-tuning on the new OmniVideo-100K dataset boosts model performance by over 20% in audio-visual reasoning tasks, revealing the power of structured scripts in enhancing multimodal understanding.
Current automated pipelines for audio-visual Question Answering (QA) generally adopt a ``video-caption-QA'' paradigm. However, these methods typically segment videos into short clips and generate separate descriptions for audio and visual modalities. This decoupled processing severs inherent associations between sounds and their visual sources, while independent clip processing often causes inconsistent descriptions of the same entity across segments. Furthermore, coupling long-text comprehension and QA synthesis into a single step often restricts models to localized events, yielding questions lacking long-term temporal connections and deep cross-modal reasoning. To address these issues, we propose an automated data engine featuring two mechanisms: (1) Entity-Anchored Video Scripting transforms videos into structured scripts, comprising summaries, main entity lists, and segment-wise audio-visual descriptions. The entity list serves as a global prior to ensure cross-segment referential consistency and reconstruct audio-visual associations. (2) Clue-Guided QA Generation prompts models to first mine cross-segment, multimodal clues from the script, and subsequently generate QA pairs based on these high-value clues. Leveraging this pipeline, we construct the instruction-tuning dataset OmniVideo-100K and a human-verified test set, OmniVideo-Test. Fine-tuning VITA-1.5, Qwen2.5-Omni-7B and Qwen3-Omni-30B on OmniVideo-100K yields performance gains of up to 20.59% on OmniVideo-Test, demonstrating strong generalization (up to 12.64% improvements) across established benchmarks like Daily-Omni and JointAVBench.