How Attentive are Graph Attention Networks?
arXiv:2105.14491
Abstract
Graph Attention Networks (GATs) are one of the most popular GNN architectures and are considered as the state-of-the-art architecture for representation learning with graphs. In GAT, every node attends to its neighbors given its own representation as the query. However, in this paper we show that GAT computes a very limited kind of attention: the ranking of the attention scores is unconditioned on the query node. We formally define this restricted kind of attention as static attention and distinguish it from a strictly more expressive dynamic attention. Because GATs use a static attention mechanism, there are simple graph problems that GAT cannot express: in a controlled problem, we show that static attention hinders GAT from even fitting the training data. To remove this limitation, we introduce a simple fix by modifying the order of operations and propose GATv2: a dynamic graph attention variant that is strictly more expressive than GAT. We perform an extensive evaluation and show that GATv2 outperforms GAT across 11 OGB and other benchmarks while we match their parametric costs. Our code is available at https://github.com/tech-srl/how_attentive_are_gats . GATv2 is available as part of the PyTorch Geometric library, the Deep Graph Library, and the TensorFlow GNN library.
Published in ICLR 2022
References in corpus (10)
- Fast Graph Representation Learning with PyTorch Geometric
- Simplifying Graph Convolutional Networks
- Interaction Networks for Learning about Objects, Relations and Physics
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- A Generalization of Transformer Networks to Graphs
- Relational Graph Attention Networks
- How to Find Your Friendly Neighborhood: Graph Attention Design with Self-Supervision
- Combining Label Propagation and Simple Models Out-performs Graph Neural Networks
- Bag of Tricks for Node Classification with Graph Neural Networks
- Improving Graph Attention Networks with Large Margin-based Constraints
Cited by in corpus (4)
- Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal Attention
- Finite-difference-informed graph network for solving steady-state incompressible flows on block-structured grids
- Enel: Context-Aware Dynamic Scaling of Distributed Dataflow Jobs using Graph Propagation
- FDGATII : Fast Dynamic Graph Attention with Initial Residual and Identity Mapping