Understanding and Resolving Performance Degradation in Graph Convolutional Networks
arXiv:2006.07107
Abstract
A Graph Convolutional Network (GCN) stacks several layers and in each layer performs a PROPagation operation (PROP) and a TRANsformation operation (TRAN) for learning node representations over graph-structured data. Though powerful, GCNs tend to suffer performance drop when the model gets deep. Previous works focus on PROPs to study and mitigate this issue, but the role of TRANs is barely investigated. In this work, we study performance degradation of GCNs by experimentally examining how stacking only TRANs or PROPs works. We find that TRANs contribute significantly, or even more than PROPs, to declining performance, and moreover that they tend to amplify node-wise feature variance in GCNs, causing variance inflammation that we identify as a key factor for causing performance drop. Motivated by such observations, we propose a variance-controlling technique termed Node Normalization (NodeNorm), which scales each node's features using its own standard deviation. Experimental results validate the effectiveness of NodeNorm on addressing performance degradation of GCNs. Specifically, it enables deep GCNs to outperform shallow ones in cases where deep models are needed, and to achieve comparable results with shallow ones on 6 benchmark datasets. NodeNorm is a generic plug-in and can well generalize to other GNN architectures. Code is publicly available at https://github.com/miafei/NodeNorm.
CIKM 2021
References in corpus (12)
- Semi-Supervised Classification with Graph Convolutional Networks
- Fast Graph Representation Learning with PyTorch Geometric
- Simplifying Graph Convolutional Networks
- Variational Graph Auto-Encoders
- Open Graph Benchmark: Datasets for Machine Learning on Graphs
- DeeperGCN: All You Need to Train Deeper GCNs
- Geom-GCN: Geometric Graph Convolutional Networks
- Combining Label Propagation and Simple Models Out-performs Graph Neural Networks
- Residual Correlation in Graph Neural Network Regression
- Revisiting Over-smoothing in Deep GCNs
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- Wiki-CS: A Wikipedia-Based Benchmark for Graph Neural Networks
Cited by in corpus (10)
- Dirichlet Energy Constrained Learning for Deep Graph Neural Networks
- Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks
- Embedding API Dependency Graph for Neural Code Generation
- Evaluating Deep Graph Neural Networks
- Residual Network and Embedding Usage: New Tricks of Node Classification with Graph Convolutional Networks
- Cold Brew: Distilling Graph Node Representations with Incomplete or Missing Neighborhoods
- A pipeline for fair comparison of graph neural networks in node classification tasks
- Search For Deep Graph Neural Networks
- Multi-Level Attention Pooling for Graph Neural Networks: Unifying Graph Representations with Multiple Localities
- Does your graph need a confidence boost? Convergent boosted smoothing on graphs with tabular node features