Do Transformers Really Perform Bad for Graph Representation?
arXiv:2106.05234
Abstract
The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer.
References in corpus (14)
- Semi-Supervised Classification with Graph Convolutional Networks
- Linformer: Self-Attention with Linear Complexity
- Graph Transformer Networks
- Conformer: Convolution-augmented Transformer for Speech Recognition
- A Generalization of Transformer Networks to Graphs
- DeeperGCN: All You Need to Train Deeper GCNs
- Graph-Bert: Only Attention is Needed for Learning Graph Representations
- OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs
- Position-aware Graph Neural Networks
- Learning Graph-Level Representation for Drug Discovery
- Rethinking Graph Transformers with Spectral Attention
- Graph Warp Module: an Auxiliary Module for Boosting the Power of Graph Neural Networks in Molecular Graph Analysis
- LazyFormer: Self Attention with Lazy Update
- Breaking the Expressive Bottlenecks of Graph Neural Networks
Cited by in corpus (11)
- Knowledge Matters: Radiology Report Generation with General and Specific Knowledge
- Temporal and Heterogeneous Graph Neural Network for Financial Time Series Prediction
- Global Self-Attention as a Replacement for Graph Convolution
- Graph Neural Networks with Learnable Structural and Positional Representations
- Logiformer: A Two-Branch Graph Transformer Network for Interpretable Logical Reasoning
- Meta-Weight Graph Neural Network: Push the Limits Beyond Global Homophily
- Gophormer: Ego-Graph Transformer for Node Classification
- Improving Automatic Parallel Training via Balanced Memory Workload Optimization
- DyFormer: A Scalable Dynamic Graph Transformer with Provable Benefits on Generalization Ability
- Permutation invariant graph-to-sequence model for template-free retrosynthesis and reaction prediction
- RaWaNet: Enriching Graph Neural Network Input via Random Walks on Graphs