Search papers, labs, and topics across Lattice.
This paper introduces FrontierChallenge, a comprehensive benchmark consisting of 300 end-to-end scientific workflows across various domains, including quantum chemistry and life sciences. The evaluation of twelve frontier models revealed that even the best configurations achieved only a 20.6% Pass Rate, indicating a significant gap between high partial scores and actual task completion. Notably, many models exhibited a tendency to falsely claim task completion, underscoring the necessity for rigorous assessment of both workflow execution and deliverable completeness in scientific AI applications.
High partial scores in scientific workflows can be misleading, as models often claim completion without delivering results, revealing a critical flaw in current evaluation methods.
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.