SSN: Learning Sparse Switchable Normalization via SparsestMax
arXiv:1903.03793
Abstract
Normalization methods improve both optimization and generalization of ConvNets. To further boost performance, the recently-proposed switchable normalization (SN) provides a new perspective for deep learning: it learns to select different normalizers for different convolution layers of a ConvNet. However, SN uses softmax function to learn importance ratios to combine normalizers, leading to redundant computations compared to a single normalizer. This work addresses this issue by presenting Sparse Switchable Normalization (SSN) where the importance ratios are constrained to be sparse. Unlike and constraints that impose difficulties in optimization, we turn this constrained optimization problem into feed-forward computation by proposing SparsestMax, which is a sparse version of softmax. SSN has several appealing properties. (1) It inherits all benefits from SN such as applicability in various tasks and robustness to a wide range of batch sizes. (2) It is guaranteed to select only one normalizer for each normalization layer, avoiding redundant computations. (3) SSN can be transferred to various tasks in an end-to-end manner. Extensive experiments show that SSN outperforms its counterparts on various challenging benchmarks such as ImageNet, Cityscapes, ADE20K, and Kinetics.
10 pages, 6 figures, accepted to CVPR 2019
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Spectral Normalization for Generative Adversarial Networks
- Categorical Reparameterization with Gumbel-Softmax
- Instance Normalization: The Missing Ingredient for Fast Stylization
- The Kinetics Human Action Video Dataset
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Densely Connected Convolutional Networks
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- Group Sparse Regularization for Deep Neural Networks
- From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
- Bayesian Uncertainty Estimation for Batch Normalized Deep Networks
- Differentiable Learning-to-Normalize via Switchable Normalization
Cited by in corpus (6)
- Differentiable Learning-to-Normalize via Switchable Normalization
- Cross-Iteration Batch Normalization
- Trilevel Neural Architecture Search for Efficient Single Image Super-Resolution
- Rethinking Normalization and Elimination Singularity in Neural Networks
- Positional Normalization
- Scale Calibrated Training: Improving Generalization of Deep Networks via Scale-Specific Normalization