Search papers, labs, and topics across Lattice.
To evaluate vision-language reasoning under localized distribution shifts, the authors construct BanglaMemeX, a benchmark of 3,000 Bangla internet memes annotated across five pragmatic dimensions alongside fine-grained, human-written explanations of visual and textual metaphors. Internet memes present a critical challenge for multimodal systems because their semantics rely on implicit socio-cultural knowledge, code-mixing, and visual irony rather than literal scene perception. Benchmarking current frontier VLMs demonstrates that while models capture coarse surface-level sentiment, they struggle significantly to articulate and ground the underlying cultural metaphors.
Modern VLMs can often classify meme sentiment on the surface, but they fail dramatically when forced to decipher and explain the localized visual metaphors driving low-resource cultural content.
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.