Search papers, labs, and topics across Lattice.
This study investigates the limitations of atom-centered structural descriptors used in machine learning for atomic-scale modeling, revealing that traditional methods based on lower-order neighbor clusters can lead to descriptor degeneracies. By employing large language models, the authors demonstrate that these degeneracies can be resolved by incorporating larger clusters of neighbors, which allows for a more accurate representation of atomic structures. The findings highlight the potential for AI to bridge knowledge across scientific domains, facilitating the discovery of novel insights and methodologies.
Large language models can uncover hidden relationships in atomic structures, revealing that even seven-neighbor clusters can yield indistinguishable descriptors.
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies''reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.