Search papers, labs, and topics across Lattice.
This report introduces a training-free framework for audio-guided video object segmentation that combines Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. By decomposing the segmentation task into distinct stages and selecting appropriate foundation models for each, the approach effectively utilizes the multimodal reasoning capabilities of MLLMs to establish text-visual correspondence while generating accurate object masks. The framework achieved competitive results in the MeViS-Audio Track of the 8th LSVOS Challenge, highlighting its potential for practical applications in video analysis without the need for additional training.
Leveraging MLLMs without any training, this framework achieves competitive audio-guided video segmentation, showcasing the power of foundation models in practical applications.
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.