Search papers, labs, and topics across Lattice.
This paper introduces the Sci-ImageMiner benchmark dataset and competition, aimed at enhancing multimodal AI's ability to comprehend and reason about scientific figures in the context of atomic layer deposition and etching. Despite the strong performance of state-of-the-art models in classification and summarization, significant challenges remain in data extraction and scientific reasoning, particularly within visual question-answering tasks. The findings underscore the limitations of current approaches and highlight the need for improved domain-aware multimodal systems to advance scientific figure comprehension.
State-of-the-art multimodal models excel in classification but falter in extracting critical data from scientific figures, revealing a significant gap in AI's reasoning capabilities.
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.