Surrogate- and invariance-boosted contrastive learning for data-scarce applications in science
arXiv:2110.08406 · doi:10.1038/s41467-022-31915-y
Abstract
Deep learning techniques have been increasingly applied to the natural sciences, e.g., for property prediction and optimization or material discovery. A fundamental ingredient of such approaches is the vast quantity of labelled data needed to train the model; this poses severe challenges in data-scarce settings where obtaining labels requires substantial computational or labor resources. Here, we introduce surrogate- and invariance-boosted contrastive learning (SIB-CL), a deep learning framework which incorporates three ``inexpensive'' and easily obtainable auxiliary information sources to overcome data scarcity. Specifically, these are: 1)~abundant unlabeled data, 2)~prior knowledge of symmetries or invariances and 3)~surrogate data obtained at near-zero cost. We demonstrate SIB-CL's effectiveness and generality on various scientific problems, e.g., predicting the density-of-states of 2D photonic crystals and solving the 3D time-independent Schrodinger equation. SIB-CL consistently results in orders of magnitude reduction in the number of labels needed to achieve the same network accuracies.
21 pages, 10 figures
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Bootstrap your own latent: A new approach to self-supervised Learning
- Molecular Contrastive Learning of Representations via Graph Neural Networks
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- Metallo-dielectric diamond and zinc-blende photonic crystals
- Overcoming data scarcity with transfer learning
- Augmentations in Hypergraph Contrastive Learning: Fabricated and Generative