Search papers, labs, and topics across Lattice.
This paper explores the application of ensemble consensus in semi-supervised learning for molecular graphs, addressing the challenge of acquiring labeled data in molecular sciences. The authors demonstrate that this approach enhances predictive accuracy and model robustness across various molecular datasets and graph neural network architectures. Notably, an individual model trained with ensemble consensus outperforms traditional ensemble methods, while also reducing calibration error, highlighting a significant advancement in semi-supervised learning techniques for molecular applications.
Ensemble consensus training not only boosts predictive accuracy in molecular graphs but also allows individual models to outperform traditional ensembles, reshaping our approach to semi-supervised learning in this domain.
Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.