Search papers, labs, and topics across Lattice.
This study investigates how multimodal large language models (MLLMs) interpret visualizations by analyzing 102 visualizations across various conditions, including access to images and contextual information. The findings reveal that providing accessible chart context significantly influences the models' tendency to make DIRECT claims and enhances numeric agreement in some cases, while the inclusion of images does not consistently improve outcomes. The research highlights the need for systems that clearly differentiate between claims grounded in evidence and those based on model-generated interpretations.
Accessible context can shift MLLMs toward more reliable claims, but images alone may not enhance interpretative accuracy.
Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, source-specific accessible chart context, and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation