Search papers, labs, and topics across Lattice.
This paper introduces onepot-Bench 0, a new benchmark suite designed to evaluate the capabilities of language models in synthetic chemistry tasks relevant to laboratory execution. The benchmark comprises three distinct evaluations: ChemAbacus for cheminformatics literacy, SynthRefusal for assessing safety and refusal behaviors, and SynthBench for predicting reaction outcomes and catalyst selection using proprietary experimental data. The findings highlight that existing evaluations often fail to capture the nuanced decision-making skills required for reliable laboratory performance, underscoring the need for more targeted assessment tools in AI-driven scientific research.
Language models struggle with lab-relevant tasks, but onepot-Bench 0 reveals critical gaps in their decision-making abilities that could impact real-world applications.
Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.