Search papers, labs, and topics across Lattice.
The Large Discovery Model (LDM) integrates a generative model with a Bayesian non-parametric reward surrogate to optimize scientific discovery across diverse domains such as neural-network training and molecular optimization. By continuously updating its discovery memory and surrogate model with new experimental data, LDM provides uncertainty-aware evaluations that significantly enhance candidate generation and selection. In empirical evaluations, LDM outperformed traditional methods, achieving a 2.4x reduction in validation BPB and over 60% gains in multi-objective performance, demonstrating its potential as a versatile discovery engine.
LDM achieves over 60% gains in multi-objective performance, revolutionizing how we approach open-ended hypothesis spaces in scientific discovery.
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a $2.4\times$ greater reduction in validation BPB, an $18.2\%$ relative decrease in binding energy, and more than $60\%$ relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.