DyFormer: A Scalable Dynamic Graph Transformer with Provable Benefits on Generalization Ability
arXiv:2111.10447
Abstract
Transformers have achieved great success in several domains, including Natural Language Processing and Computer Vision. However, its application to real-world graphs is less explored, mainly due to its high computation cost and its poor generalizability caused by the lack of enough training data in the graph domain. To fill in this gap, we propose a scalable Transformer-like dynamic graph learning method named Dynamic Graph Transformer (DyFormer) with spatial-temporal encoding to effectively learn graph topology and capture implicit links. To achieve efficient and scalable training, we propose temporal-union graph structure and its associated subgraph-based node sampling strategy. To improve the generalization ability, we introduce two complementary self-supervised pre-training tasks and show that jointly optimizing the two pre-training tasks results in a smaller Bayesian error rate via an information-theoretic analysis. Extensive experiments on the real-world datasets illustrate that DyFormer achieves a consistent 1%-3% AUC gain (averaged over all time steps) compared with baselines on all benchmarks.
References in corpus (11)
- KGAT: Knowledge Graph Attention Network for Recommendation
- Graph Convolutional Matrix Completion
- Linformer: Self-Attention with Linear Complexity
- Graph Transformer Networks
- Exploring Simple Siamese Representation Learning
- A Generalization of Transformer Networks to Graphs
- Graph-Bert: Only Attention is Needed for Learning Graph Representations
- Do Transformers Really Perform Bad for Graph Representation?
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Inductive Representation Learning on Temporal Graphs
- Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks