Shake-Shake regularization
arXiv:1705.07485
Abstract
The method introduced in this paper aims at helping deep learning practitioners faced with an overfit problem. The idea is to replace, in a multi-branch network, the standard summation of parallel branches with a stochastic affine combination. Applied to 3-branch residual networks, shake-shake regularization improves on the best single shot published results on CIFAR-10 and CIFAR-100 by reaching test errors of 2.86% and 15.85%. Experiments on architectures without skip connections or Batch Normalization show encouraging results and open the door to a large set of applications. Code is available at https://github.com/xgastaldi/shake-shake
Cited by in corpus (13)
- Improved Regularization of Convolutional Neural Networks with Cutout
- A Survey on Neural Architecture Search
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- Weighted Transformer Network for Machine Translation
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- Neural SDE: Stabilizing Neural ODE Networks with Stochastic Noise
- Confidence Calibration for Convolutional Neural Networks Using Structured Dropout
- Connectivity Learning in Multi-Branch Networks
- Parle: parallelizing stochastic gradient descent
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Genetic Network Architecture Search
- Variations on the Chebyshev-Lagrange Activation Function
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification