On Representation Knowledge Distillation for Graph Neural Networks
arXiv:2111.04964 · doi:10.1109/TNNLS.2022.3223018
Abstract
Knowledge distillation is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the Local Structure Preserving loss (LSP), which matches local structural relationships defined over edges across the student and teacher's node embeddings. This paper studies whether preserving the global topology of how the teacher embeds graph data can be a more effective distillation objective for GNNs, as real-world graphs often contain latent interactions and noisy edges. We propose Graph Contrastive Representation Distillation (G-CRD), which uses contrastive learning to implicitly preserve global topology by aligning the student node embeddings to those of the teacher in a shared representation space. Additionally, we introduce an expanded set of benchmarks on large-scale real-world datasets where the performance gap between teacher and student GNNs is non-negligible. Experiments across 4 datasets and 14 heterogeneous GNN architectures show that G-CRD consistently boosts the performance and robustness of lightweight GNNs, outperforming LSP (and a global structure preserving variant of LSP) as well as baselines from 2D computer vision. An analysis of the representational similarity among teacher and student embedding spaces reveals that G-CRD balances preserving local and global relationships, while structure preserving approaches are best at preserving one or the other. Our code is available at https://github.com/chaitjo/efficient-gnns
IEEE Transactions on Neural Networks and Learning Representation (TNNLS), Special Issue on Deep Neural Networks for Graphs: Theory, Models, Algorithms and Applications
References in corpus (22)
- Distilling the Knowledge in a Neural Network
- Knowledge Distillation: A Survey
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Fast Graph Representation Learning with PyTorch Geometric
- GCC: Graph Contrastive Coding for Graph Neural Network Pre-Training
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- Similarity of Neural Network Representations Revisited
- Benchmarking Graph Neural Networks
- Fake News Detection on Social Media using Geometric Deep Learning
- SIGN: Scalable Inception Graph Neural Networks
- Pre-training Molecular Graph Representation with 3D Geometry
- Combining Label Propagation and Simple Models Out-performs Graph Neural Networks
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- 3D Infomax improves GNNs for Molecular Property Prediction
- Rethinking pooling in graph neural networks
- Graph-less Neural Networks: Teaching Old MLPs New Tricks via Distillation
- Probabilistic Dual Network Architecture Search on Graphs
- Learned Low Precision Graph Neural Networks
- Large-scale graph representation learning with very deep GNNs and self-supervision
- Iterative Graph Self-Distillation
- Training Graph Neural Networks with 1000 Layers