Search papers, labs, and topics across Lattice.
This paper introduces PragMatch, a benchmark designed to assess the ability of Large Vision-Language Models (LVLMs) to distinguish between genuine pragmatic incongruity and superficial correlations in multimodal contexts, particularly for sarcasm detection. Through the analysis of 3,000 curated image-text pairs and targeted injection experiments, the authors reveal that LVLMs are significantly influenced by lexical and stylistic cues, leading to substantial shifts in predictions even when the underlying relationships remain unchanged. The findings highlight critical limitations in current LVLMs and underscore the need for more robust evaluation methods for multimodal reasoning capabilities.
LVLMs are misled by superficial cues, with injected signals causing drastic shifts in sarcasm detection accuracy, revealing a fundamental flaw in their reasoning abilities.
Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.