Search papers, labs, and topics across Lattice.
This paper tackles the challenge of determining molecular structures from spectroscopic data by introducing a scalable hypothesis-refinement paradigm that combines spectral evidence with extensive molecular priors. The authors construct the QM9SPIN dataset, which includes diverse NMR spectra, and develop two models: SpectroMol for generating chemically valid molecular hypotheses and MS-Mol2Mol for refining these structures using mass constraints. Their integrated system achieves a remarkable 93.8% top-1 accuracy on a simulated benchmark and shows effective adaptation to experimental data, paving the way for automated organic structure elucidation.
Achieving 93.8% accuracy in predicting molecular structures from multimodal spectroscopic data could revolutionize how chemists approach organic synthesis and analysis.
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.