Search papers, labs, and topics across Lattice.
This paper introduces NARU, a benchmark specifically designed to assess narrative evolution and cultural nuance understanding in Japanese long-form videos, addressing a significant gap in existing evaluation frameworks. The benchmark comprises 1,481 questions derived from 155 videos, employing a hierarchical memory-based annotation pipeline that ensures thorough event, narrative, and cultural context representation. Evaluation results highlight considerable deficiencies in current models' abilities to integrate long-range narratives and perform culturally informed reasoning, underscoring the need for improved methodologies in multi-modal language model training.
Current models struggle with long-range narrative integration and cultural reasoning, revealing critical gaps in their understanding of high-context media.
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.