Shake-Shake regularization
arXiv:1705.07485
Abstract
The method introduced in this paper aims at helping deep learning practitioners faced with an overfit problem. The idea is to replace, in a multi-branch network, the standard summation of parallel branches with a stochastic affine combination. Applied to 3-branch residual networks, shake-shake regularization improves on the best single shot published results on CIFAR-10 and CIFAR-100 by reaching test errors of 2.86% and 15.85%. Experiments on architectures without skip connections or Batch Normalization show encouraging results and open the door to a large set of applications. Code is available at https://github.com/xgastaldi/shake-shake
Cited by in corpus (54)
- Improved Regularization of Convolutional Neural Networks with Cutout
- A Survey on Neural Architecture Search
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- Weighted Transformer Network for Machine Translation
- Selective Kernel Networks
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- Neural SDE: Stabilizing Neural ODE Networks with Stochastic Noise
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning
- SELF: Learning to Filter Noisy Labels with Self-Ensembling
- Confidence Calibration for Convolutional Neural Networks Using Structured Dropout
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- sharpDARTS: Faster and More Accurate Differentiable Architecture Search
- Optimized Generic Feature Learning for Few-shot Classification across Domains
- Connectivity Learning in Multi-Branch Networks
- Efficient Certified Defenses Against Patch Attacks on Image Classifiers
- Parle: parallelizing stochastic gradient descent
- Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need
- Robust Learning Under Label Noise With Iterative Noise-Filtering
- Semi-Supervised Learning by Label Gradient Alignment
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Group Ensemble: Learning an Ensemble of ConvNets in a single ConvNet
- KeepAugment: A Simple Information-Preserving Data Augmentation Approach
- Improving Neural Architecture Search Image Classifiers via Ensemble Learning
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- A Novel Framework for Neural Architecture Search in the Hill Climbing Domain
- CacheNet: A Model Caching Framework for Deep Learning Inference on the Edge
- Evaluating State-of-the-Art Classification Models Against Bayes Optimality
- On the Effectiveness of Regularization Against Membership Inference Attacks
- AdaSGD: Bridging the gap between SGD and Adam
- Bag of Tricks for Neural Architecture Search
- Learning Optimal Data Augmentation Policies via Bayesian Optimization for Image Classification Tasks
- Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers
- OnlineAugment: Online Data Augmentation with Less Domain Knowledge
- Effective Evaluation of Deep Active Learning on Image Classification Tasks
- Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition
- Image recognition from raw labels collected without annotators
- Fine-tuning Handwriting Recognition systems with Temporal Dropout
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective
- Genetic Network Architecture Search
- Multi-level Feature Fusion-based CNN for Local Climate Zone Classification from Sentinel-2 Images: Benchmark Results on the So2Sat LCZ42 Dataset
- MetaAugment: Sample-Aware Data Augmentation Policy Learning
- Variations on the Chebyshev-Lagrange Activation Function
- Meta Gradient Adversarial Attack
- Hierarchical Auxiliary Learning
- FROST: Faster and more Robust One-shot Semi-supervised Training
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- Improve SGD Training via Aligning Mini-batches
- RotationOut as a Regularization Method for Neural Network
- A Comprehensive Study on Optimization Strategies for Gradient Descent In Deep Learning
- Interlayer and Intralayer Scale Aggregation for Scale-invariant Crowd Counting
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum