Search papers, labs, and topics across Lattice.
This study leverages multimodal literature mining to convert inaccessible X-ray absorption spectroscopy (XAS) data embedded in figures and text into a structured, AI-ready dataset. By developing a scalable digitization pipeline, the authors extracted 13,740 XAS spectra from battery literature, covering 66 absorbing elements and various chemistries, with expert validation confirming the accuracy of the data. This transformation not only enhances the accessibility of XAS data but also lays the groundwork for advanced material discovery and high-throughput analysis across laboratories.
Transforming fragmented XAS data into a structured dataset unlocks unprecedented opportunities for high-throughput material discovery in battery research.
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.