Search papers, labs, and topics across Lattice.
This paper introduces MultivationBench, a benchmark aimed at evaluating multimodal sequential motivation reasoning in Large Language Models (LLMs) through story-driven visual narratives. By leveraging psychological frameworks such as Maslow's hierarchy and Reiss's basic desires, the benchmark assesses models' abilities to integrate and reason about accumulated multimodal context over time. The findings reveal that current models face substantial challenges in maintaining consistent motivation reasoning across sequential contexts, highlighting a gap between their static recognition abilities and the dynamic reasoning needed for human-like social intelligence.
Current multimodal models falter in maintaining consistent motivation reasoning across sequences, exposing a critical gap in their social intelligence capabilities.
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.