Search papers, labs, and topics across Lattice.
This paper introduces MarineEVT, the first event-centric dataset for marine video understanding, comprising 20,000 multi-task video-level visual question-answering pairs. The authors propose EVT-R1, a novel reasoning framework that integrates visual tools to enhance the localization and interpretation of critical information in marine videos, addressing the challenges posed by sparse and unpredictable events. Experimental results show that EVT-R1 significantly outperforms leading vision-language models, achieving improvements of 5.22 and 11.09 over top open-source and commercial models, respectively.
Marine video understanding can be revolutionized with a new dataset and reasoning framework that outperforms existing models by over 11 points.
Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.