Search papers, labs, and topics across Lattice.
This paper critiques the limitations of Large Vision-Language Models (LVLMs) in evaluating the temporal coherence of visual narratives, highlighting a significant performance gap in their ability to discriminate between coherent and incoherent sequences. The authors reveal that LVLMs are influenced by positional biases, such as primacy and recency effects, which impair their judgment of narrative flow. By identifying these structural flaws, the study calls for a shift towards Temporally-Aware Evaluation paradigms that better reflect the logical continuity of visual storytelling.
LVLMs struggle with temporal reasoning, showing that their judgment can be swayed more by frame placement than by narrative coherence itself.
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.