Search papers, labs, and topics across Lattice.
This paper introduces PertReason, a benchmark and framework designed to evaluate mechanistic reasoning in machine learning models regarding perturbation effects in cellular contexts. By integrating single-cell perturbation data with knowledge graphs and conditioning on cell-specific states, PertReasonQA assesses whether models can provide accurate explanations while remaining robust to distribution shifts. The findings reveal significant discrepancies between predictive accuracy and mechanistic reasoning, highlighting common failure modes in state-of-the-art models that are not captured by traditional benchmarks.
Models that excel in prediction often falter in mechanistic reasoning, revealing hidden flaws in their logic and understanding of cellular context.
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for cell-state--conditioned reasoning about perturbation effects. At its core, PertReasonQA is a benchmark that tests whether models can generate mechanistically faithful explanations while remaining robust to complex shifts, such as new cells and unseen perturbations. PertReasonQA combines single-cell genetic and chemical perturbation data across multiple cellular contexts with knowledge graphs, and dynamically conditions pathways on cell-specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PertReasonLM, a large language model trained to align outcome predictions with context-specific mechanistic reasoning. Our model targets the identified failure modes by grounding rationales in context-specific pathways and tightening agreement between outcomes and mechanisms. Together, we provide a diagnostic framework for exposing and mitigating failures in faithful reasoning in data-rich scientific systems.