NewEvery arXiv paper, its researchers & institutions — mapped.
machine learning

Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus

arXiv:2607.28304

summary

The paper proposes a semi-supervised learning approach for molecular graph data that uses an ensemble consensus objective to improve prediction accuracy, robustness, and calibration across various datasets and GNN architectures.

Abstract

Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.

ICML

Topics & keywords

#semi-supervised learning#molecular graphs#graph neural networks#ensemble methods#knowledge distillationensemble consensusgraph neural networkcalibration errorproperty predictionsemi-supervised training