Search papers, labs, and topics across Lattice.
This study explores the use of open-weight large language models (LLMs) as acquisition policies for optimizing materials from finite candidate pools, addressing the high costs associated with experimental evaluations. By evaluating five LLMs across four materials optimization tasks, the researchers found that LLMs can reach global optima in fewer iterations than random selection, although their performance compared to conventional Gaussian-process methods varies significantly. The findings highlight the potential of LLMs to provide valuable acquisition signals without the need for task-specific training, though their effectiveness is influenced by task complexity and candidate presentation strategies.
LLMs can outperform random selection in materials optimization, but their effectiveness varies widely across tasks and contexts.
Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this setting. We evaluate five LLMs across four retrospective finite-pool materials optimization tasks under different candidate-presentation strategies and compare them with random selection and conventional Gaussian-process methods. LLM policies generally reach the global optimum in fewer iterations than random selection, indicating that they provide a useful acquisition signal without task-specific training. Their performance relative to Gaussian-process methods is mixed: conventional acquisition performs better on most tasks, while LLMs match or outperform it in some settings. Performance varies substantially across tasks, models, initializations, and candidate presentations, with no LLM approach performing best across all tasks. Overall, open-weight LLMs show potential as acquisition policies for finite-pool materials search, although their reliability remains sensitive to the task and to how candidates and scientific context are presented.