Search papers, labs, and topics across Lattice.
This paper introduces LongChart, a new benchmark designed to assess the capabilities of multimodal large language models (MLLMs) in complex multi-chart reasoning tasks. By employing a synthesis pipeline supported by latent graphs, the authors create a dataset that includes an average of 6.5 images and 31.2 questions per visual question answering (VQA) set, addressing the limitations of existing benchmarks that focus on single-chart perception. The evaluation of 10 state-of-the-art MLLMs reveals significant performance degradation as computational complexity increases, underscoring the need for further research in this area.
MLLM accuracy drops significantly with increased computational complexity, revealing critical gaps in current benchmarks for multi-chart reasoning.
Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.