Search papers, labs, and topics across Lattice.
This paper introduces VideoGAIA, a novel benchmark designed to assess the agentic video understanding capabilities of multimodal large language models (MLLMs) through multi-turn, tool-augmented interactions. Unlike traditional single-turn tasks, VideoGAIA requires models to iteratively engage with video content, utilize external tools, and synthesize information across multiple interactions, reflecting real-world complexities. The results reveal that even leading models like GPT-5.5 and Kimi-K3 struggle with this benchmark, achieving less than 60% accuracy, underscoring the need for more sophisticated evaluation methods in AI development.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.