CASRuhr University BochumShenzhen Loop Area InstituteTencent AIMay 31, 2026arXiv:2606.01348

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Shangpin Peng, Gengluo Li, Xingyu Wan, Chengquan Zhang, Hao Feng, Binghong Wu, Huawen Shen, Weinong Wang, Ziyi Cai, Zhuotao Tian, Han Hu, Can Ma, Yu Zhou

AI Summary

This paper introduces ChartArena, a comprehensive bilingual benchmark designed to evaluate chart parsing models across eight chart families and three visual scenarios, including digital, printed, and hand-drawn formats. By employing a human-agent collaborative annotation pipeline and a format-agnostic evaluation protocol, the study provides a reliable framework for assessing model performance and facilitating cross-model comparisons. The evaluation of 26 leading MLLMs reveals significant capability gaps, particularly in handling diagrammatic structures and challenging scenarios like radar charts and hand-drawn images, highlighting the need for further advancements in chart parsing technologies.

Key Contribution

ChartArena reveals that even top proprietary models struggle with diagrammatic structures, exposing critical gaps in current chart parsing capabilities.

Abstract

Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult. Existing benchmarks focus on narrow chart types and leave diagrammatic structures such as flowcharts and mind maps largely unaddressed, while models produce outputs in incompatible formats, and datasets rarely include the printed or hand-drawn images encountered in practice. To address these issues, we introduce ChartArena, a comprehensive bilingual benchmark covering eight chart families spanning both numeric charts and diagrammatic structures, each evaluated across three visual scenarios: digital renderings, printed photos, and hand-drawn photos. The dataset is built via a human-agent collaborative annotation pipeline with multi-stage human verification to ensure annotation reliability. To enable fair cross-model comparison, we further design a format-agnostic evaluation protocol that maps heterogeneous outputs into two canonical semantic spaces, a normalized triple view and a directed graph view, and scores them with structure-aware metrics. Through extensive evaluation of 26 leading MLLMs, we observe three consistent findings: (i) frontier proprietary models such as Gemini 3.1 Pro lead overall, yet the strongest open-source systems are rapidly closing the gap; (ii) document parsing models handle numeric charts reasonably but fall sharply behind on diagrammatic structures; and (iii) expert chart parsers remain limited to narrow chart families. Across all models, radar charts and hand-drawn scenarios stay especially challenging. These findings show that ChartArena exposes clear capability gaps and provides a unified foundation for future progress. ChartArena is publicly available at https://github.com/pspdada/ChartArena.

Eval Frameworks & Benchmarks Multimodal Models

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Related Papers