Search papers, labs, and topics across Lattice.
This paper introduces the ALD/E-ImageMiner benchmark and the ICDAR 2026 Competition, which collectively provide a dataset of 1,951 expert-annotated scientific figures for tasks such as classification, data extraction, and visual question answering. By exploring how these tasks assess capabilities from visual comprehension to domain-specific reasoning, the authors argue for a long-term objective of achieving "scientific conceptual understanding from images." The findings suggest that advancing multimodal AI in scientific contexts can significantly enhance the retrieval and interpretation of complex visual data, paving the way for more effective scientific communication and knowledge extraction.
Unlocking the potential of multimodal AI to interpret complex scientific images could revolutionize how we access and understand experimental data.
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.