Search papers, labs, and topics across Lattice.
This paper introduces PersonaShot, a novel benchmark designed to evaluate narrative continuity in multi-shot video generation, focusing on the coherence of physical and emotional states of characters across cuts. By incorporating approximately 1,000 multi-shot segments and 16 specific metrics, the benchmark assesses continuity at various temporal levels, revealing significant gaps in state-of-the-art models' performance regarding cross-shot narrative coherence. The findings highlight that even visually impressive videos often fail to maintain character consistency, underscoring the need for improved methods in video generation that prioritize narrative integrity.
Even the most visually stunning video generation models struggle to maintain character continuity across shots, revealing a critical gap in current evaluation methods.
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.