Search papers, labs, and topics across Lattice.
This study introduces AtmosCoder-Bench, a novel execution-grounded benchmark designed to evaluate the calculation processes of large language models (LLMs) in environmental science, revealing critical insights into their performance. The research uncovers that traditional multiple-choice evaluations can inflate accuracy metrics by at least 12 percentage points and highlights that many errors stem from LLMs' inconsistent application of known formulas during multi-step calculations. Furthermore, it demonstrates that even state-of-the-art models struggle to adapt their reasoning to specific task conditions, underscoring the necessity for expert oversight in their deployment.
LLMs can miss 40% of the necessary calculations in environmental science, revealing a critical gap in their reliability for quantitative tasks.
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human-solvable, with uniquely verifiable answers. We find that (i) multiple-choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi-step computation; and (iii) even frontier models remain weak when task-specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.