Search papers, labs, and topics across Lattice.
This paper introduces DrivelHub+, a benchmark designed to evaluate video-language models on their ability to interpret implicit and non-literal meanings in social media videos, which often rely on multimodal cues and cultural context. The benchmark consists of 1,000 annotated videos, each accompanied by a human-written implicit narrative, challenging models to go beyond mere recognition and description. Evaluations focus on both the models' ability to explain pragmatic comprehension and their capacity to align video representations with corresponding narratives, revealing significant gaps in current multimodal reasoning capabilities.
Current video-language models struggle to interpret the nuanced meanings behind social media videos, often missing the implicit narratives that make them humorous or ironic.
Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.