Search papers, labs, and topics across Lattice.
This study investigates the reliability of chemical reasoning language models using chain-of-thought (CoT) approaches across various chemistry tasks, revealing that hallucinations are prevalent and often occur alongside correct answers. The authors identify a shared scratchpad function in different model families, where models like Chem-R and ether-0 utilize fragmented SMILES drafts, while ChemDFM-R focuses on structural cues. Importantly, the research demonstrates that these structural drafts play a crucial role in generation, suggesting that CoT outputs should not be viewed as definitive evidence of accurate reasoning.
Hallucinations in chemical reasoning models coexist with correct answers, revealing a complex relationship that challenges our understanding of model reliability.
Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.