Pre-training Molecular Graph Representation with 3D Geometry
arXiv:2110.07728
Abstract
Molecular graph representation learning is a fundamental problem in modern drug and material discovery. Molecular graphs are typically modeled by their 2D topological structures, but it has been recently discovered that 3D geometric information plays a more vital role in predicting molecular functionalities. However, the lack of 3D information in real-world scenarios has significantly impeded the learning of geometric graph representation. To cope with this challenge, we propose the Graph Multi-View Pre-training (GraphMVP) framework where self-supervised learning (SSL) is performed by leveraging the correspondence and consistency between 2D topological structures and 3D geometric views. GraphMVP effectively learns a 2D molecular graph encoder that is enhanced by richer and more discriminative 3D geometry. We further provide theoretical insights to justify the effectiveness of GraphMVP. Finally, comprehensive experiments show that GraphMVP can consistently outperform existing graph SSL methods.
References in corpus (13)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Graph Self-Supervised Learning: A Survey
- Contrastive Multi-View Representation Learning on Graphs
- On Variational Bounds of Mutual Information
- E(n) Equivariant Graph Neural Networks
- Graph Contrastive Learning Automated
- How to Train Your Energy-Based Models
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Spherical Message Passing for 3D Graph Networks
- Self-supervised Graph-level Representation Learning with Local and Global Structure
- Improved Contrastive Divergence Training of Energy Based Models
- Message Passing Networks for Molecules with Tetrahedral Chirality
Cited by in corpus (10)
- Improving Molecular Contrastive Learning via Faulty Negative Mitigation and Decomposed Fragment Contrast
- GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
- On Representation Knowledge Distillation for Graph Neural Networks
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property Prediction
- Automated 3D Pre-Training for Molecular Property Prediction
- A Text-guided Protein Design Framework
- Multi-channel learning for integrating structural hierarchies into context-dependent molecular representation
- A Large Encoder-Decoder Family of Foundation Models For Chemical Language
- BatmanNet: Bi-branch Masked Graph Transformer Autoencoder for Molecular Representation
- Rotation-equivariant Graph Neural Networks for Learning Glassy Liquids Representations