Search papers, labs, and topics across Lattice.
This study fine-tunes Large Language Models (LLMs) to predict molecular geometries by leveraging both Cartesian coordinates and Z-matrices, demonstrating that Z-matrix representation significantly enhances prediction accuracy. The results indicate that fine-tuned LLMs can outperform specialized deep learning models in predicting equilibrium structures and conformers of small organic and drug-like molecules. Importantly, the approach retains the LLM's pre-trained language capabilities by incorporating natural language prompts during fine-tuning, showcasing a balance between domain-specific performance and general language understanding.
Z-matrices provide a superior grammar for LLM adaptation, leading to unprecedented accuracy in predicting molecular geometries while preserving language capabilities.
The power of Large Language Models (LLMs) has led us to investigate how they might be fine-tuned for learning the"language of molecular geometry". The fine-tuning of the LLMs using Cartesian coordinates or Z-matrices provides an extremely simple method for accurately predicting equilibrium structures and diverse sets of conformers of small organic and drug-like molecules with excellent accuracy and outperforming specialized deep learning models. While the most common representation of molecular geometries are Cartesian coordinates performs adequately, we find that the inherent invariances and relational nature of geometries represented as Z-matrices provides a better grammar for LLM adaptation. Finally, we show that enhancing an LLMs capabilities for robust prediction of small molecule geometries still retains nearly all of its pre-trained language abilities by randomly mixing in small quantities of natural language prompt-response pairs into the fine-tuning.