Search papers, labs, and topics across Lattice.
This paper introduces Top-K prompting as a novel training and inference paradigm for single-step retrosynthesis, addressing the limitations of traditional single-answer evaluation methods. By leveraging an ultra-large-scale dataset of approximately 45.6 million verified reactions, the authors train the Chemistry Constraint-Consistent Language Model (C3LM) and achieve state-of-the-art performance on the OOD URSA-expert-2026 benchmark through fine-tuning with ChemCensor-based and novelty-oriented rewards. The findings reveal that LLMs and conventional models explore complementary reaction spaces, suggesting the potential for ensemble-based systems in retrosynthesis planning.
Top-K prompting transforms retrosynthesis by capturing diverse reaction predictions, outperforming traditional models in accuracy and uniqueness.
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.