Search papers, labs, and topics across Lattice.
The paper introduces ScienceArena, a benchmark designed to evaluate large language models (LLMs) using complex, open-ended problems derived from prestigious science competitions in physics, chemistry, and biology. By employing a rigorous digitization pipeline and expert-audited scoring rubrics, the authors calibrate LLMs against human medalist performance, revealing that while top models achieve impressive scores, they struggle with visual grounding and maintaining consistency over longer tasks. The findings highlight critical areas for improvement in LLMs, particularly in chemistry and problem-solving coherence, underscoring the need for enhanced reasoning capabilities in AI systems.
Top LLMs can achieve medal-equivalent scores on elite science exams, but they falter on visual grounding and long-horizon consistency.
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.