A Robust Initialization of Residual Blocks for Effective ResNet Training without Batch Normalization
arXiv:2112.12299 · doi:10.1109/TNNLS.2023.3325541
Abstract
Batch Normalization is an essential component of all state-of-the-art neural networks architectures. However, since it introduces many practical issues, much recent research has been devoted to designing normalization-free architectures. In this paper, we show that weights initialization is key to train ResNet-like normalization-free networks. In particular, we propose a slight modification to the summation operation of a block output to the skip-connection branch, so that the whole network is correctly initialized. We show that this modified architecture achieves competitive results on CIFAR-10, CIFAR-100 and ImageNet without further regularization nor algorithmic modifications.
16 pages (4 pages of supplementary material), 9 figures, 2 table
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- High-Performance Large-Scale Image Recognition Without Normalization
- Delving Deep into Label Smoothing
- ReZero is All You Need: Fast Convergence at Large Depth
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch Normalization