BatmanNet: Bi-branch Masked Graph Transformer Autoencoder for Molecular Representation
arXiv:2211.13979 · doi:10.1093/bib/bbad400
Abstract
Although substantial efforts have been made using graph neural networks (GNNs) for AI-driven drug discovery (AIDD), effective molecular representation learning remains an open challenge, especially in the case of insufficient labeled molecules. Recent studies suggest that big GNN models pre-trained by self-supervised learning on unlabeled datasets enable better transfer performance in downstream molecular property prediction tasks. However, the approaches in these studies require multiple complex self-supervised tasks and large-scale datasets, which are time-consuming, computationally expensive, and difficult to pre-train end-to-end. Here, we design a simple yet effective self-supervised strategy to simultaneously learn local and global information about molecules, and further propose a novel bi-branch masked graph transformer autoencoder (BatmanNet) to learn molecular representations. BatmanNet features two tailored complementary and asymmetric graph autoencoders to reconstruct the missing nodes and edges, respectively, from a masked molecular graph. With this design, BatmanNet can effectively capture the underlying structure and semantic information of molecules, thus improving the performance of molecular representation. BatmanNet achieves state-of-the-art results for multiple drug discovery tasks, including molecular properties prediction, drug-drug interaction, and drug-target interaction, on 13 benchmark datasets, demonstrating its great potential and superiority in molecular representation learning.
19 pages, 6 figures, Accepted by Briefings in Bioinformatics in 17-Oct-2023
References in corpus (11)
- Semi-Supervised Classification with Graph Convolutional Networks
- ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction
- SchNet: A continuous-filter convolutional neural network for modeling quantum interactions
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- Massively Multitask Networks for Drug Discovery
- SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery
- Pre-training Molecular Graph Representation with 3D Geometry
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property Prediction
- A Survey of Pretraining on Graphs: Taxonomy, Methods, and Applications
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- MGAE: Masked Autoencoders for Self-Supervised Learning on Graphs