Deep Networks with Stochastic Depth
arXiv:1603.09382
Abstract
Very deep convolutional networks with hundreds of layers have led to significant reductions in error on competitive benchmarks. Although the unmatched expressiveness of the many layers can be highly desirable at test time, training very deep networks comes with its own set of challenges. The gradients can vanish, the forward flow often diminishes, and the training time can be painfully slow. To address these problems, we propose stochastic depth, a training procedure that enables the seemingly contradictory setup to train short networks and use deep networks at test time. We start with very deep networks but during training, for each mini-batch, randomly drop a subset of layers and bypass them with the identity function. This simple approach complements the recent success of residual networks. It reduces training time substantially and improves the test error significantly on almost all data sets that we used for evaluation. With stochastic depth we can increase the depth of residual networks even beyond 1200 layers and still yield meaningful improvements in test error (4.91% on CIFAR-10).
first two authors contributed equally
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Striving for Simplicity: The All Convolutional Net
- Scalable Bayesian Optimization Using Deep Neural Networks
- Learning Activation Functions to Improve Deep Neural Networks
- Fractional Max-Pooling
- Identity Mappings in Deep Residual Networks
- Generalizing Pooling Functions in Convolutional Neural Networks: Mixed, Gated, and Tree
Cited by in corpus (152)
- Neural Architecture Search with Reinforcement Learning
- Densely Connected Convolutional Networks
- Wide Residual Networks
- On Calibration of Modern Neural Networks
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Temporal Ensembling for Semi-Supervised Learning
- Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
- Activation Functions: Comparison of trends in Practice and Research for Deep Learning
- ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Pruning Filters for Efficient ConvNets
- FractalNet: Ultra-Deep Neural Networks without Residuals
- Residual Networks Behave Like Ensembles of Relatively Shallow Networks
- Large-Scale Evolution of Image Classifiers
- Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
- Group Sparse Regularization for Deep Neural Networks
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Residual Networks of Residual Networks: Multilevel Residual Networks
- QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Knowledge Distillation by On-the-Fly Native Ensemble
- Shake-Shake regularization
- Residual Attention Network for Image Classification
- Learning Efficient Convolutional Networks through Network Slimming
- Hierarchical Multiscale Recurrent Neural Networks
- Dimensionality-Driven Learning with Noisy Labels
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks
- Swapout: Learning an ensemble of deep architectures
- Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- SaliencyMix: A Saliency Guided Data Augmentation Strategy for Better Regularization
- The Power of Sparsity in Convolutional Neural Networks
- Regularizing CNNs with Locally Constrained Decorrelations
- Artificial Intelligence and its Role in Near Future
- Interleaved Group Convolutions for Deep Neural Networks
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- FreezeOut: Accelerate Training by Progressively Freezing Layers
- Reversible Architectures for Arbitrarily Deep Residual Neural Networks
- Multi-level Residual Networks from Dynamical Systems View
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
- Analysis and Optimization of Convolutional Neural Network Architectures
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Multi-Sample Dropout for Accelerated Training and Better Generalization
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- Deep Convolutional Neural Network Design Patterns
- Survey of Dropout Methods for Deep Neural Networks
- Regularization and Optimization strategies in Deep Convolutional Neural Network
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Optimization on Submanifolds of Convolution Kernels in CNNs
- Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
- Patch-based Progressive 3D Point Set Upsampling
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- PolyNet: A Pursuit of Structural Diversity in Very Deep Networks
- Learning Deep Morphological Networks with Neural Architecture Search
- Pixel-wise Attentional Gating for Parsimonious Pixel Labeling
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- A Novel Weight-Shared Multi-Stage CNN for Scale Robustness
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Deep Pyramidal Residual Networks
- Log-DenseNet: How to Sparsify a DenseNet
- A Survey of Deep Learning Techniques for Mobile Robot Applications
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Training DNNs with Hybrid Block Floating Point
- Spatially Adaptive Computation Time for Residual Networks
- Crafting GBD-Net for Object Detection
- DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource Constraints
- Image Super-Resolution via Dual-State Recurrent Networks
- Ranking to Learn and Learning to Rank: On the Role of Ranking in Pattern Recognition Applications
- Deep Pyramidal Residual Networks with Separated Stochastic Depth
- ShaResNet: reducing residual network parameter number by sharing weights
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Data Dropout: Optimizing Training Data for Convolutional Neural Networks
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Rethinking Channel Dimensions for Efficient Model Design
- Projection Based Weight Normalization for Deep Neural Networks
- A Novel Convolutional Neural Network for Image Steganalysis with Shared Normalization
- Riemannian approach to batch normalization
- Bayesian Sparsification of Recurrent Neural Networks
- Dynamic Optimization of Neural Network Structures Using Probabilistic Modeling
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Efficient Stochastic Inference of Bitwise Deep Neural Networks
- Convolutional Residual Memory Networks
- Implementation of Deep Convolutional Neural Network in Multi-class Categorical Image Classification
- A Way out of the Odyssey: Analyzing and Combining Recent Insights for LSTMs
- Shifting Mean Activation Towards Zero with Bipolar Activation Functions
- Learning Better Internal Structure of Words for Sequence Labeling
- BERT's output layer recognizes all hidden layers? Some Intriguing Phenomena and a simple way to boost BERT
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Every Node Counts: Self-Ensembling Graph Convolutional Networks for Semi-Supervised Learning
- HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass Lens
- Deep Shape from Polarization
- Batch-normalized Recurrent Highway Networks
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Learning Normalized Inputs for Iterative Estimation in Medical Image Segmentation
- Approximate Dynamic Programming with Neural Networks in Linear Discrete Action Spaces
- Semi-Supervised Noisy Student Pre-training on EfficientNet Architectures for Plant Pathology Classification
- Identifying Most Walkable Direction for Navigation in an Outdoor Environment
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- Compressing Language Models using Doped Kronecker Products
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- Simplified Stochastic Feedforward Neural Networks
- Layer Folding: Neural Network Depth Reduction using Activation Linearization
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- PhytNet -- Tailored Convolutional Neural Networks for Custom Botanical Data
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- TandemNet: Distilling Knowledge from Medical Images Using Diagnostic Reports as Optional Semantic References
- Improving training of deep neural networks via Singular Value Bounding
- Towards Adversarial Training with Moderate Performance Improvement for Neural Network Classification
- Extrapolation for Large-batch Training in Deep Learning
- Deep Learning Works in Practice. But Does it Work in Theory?
- Shortcut Sequence Tagging
- GM-Net: Learning Features with More Efficiency
- On Residual Networks Learning a Perturbation from Identity
- Learning Light-Weight Translation Models from Deep Transformer
- Truncating Wide Networks using Binary Tree Architectures
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- SuperNet -- An efficient method of neural networks ensembling
- Sharpen Focus: Learning with Attention Separability and Consistency
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- Convergence of backpropagation with momentum for network architectures with skip connections
- Approximate Fisher Information Matrix to Characterise the Training of Deep Neural Networks
- Orthogonal and Idempotent Transformations for Learning Deep Neural Networks
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Frequency Disentangled Residual Network
- Reconciling Feature-Reuse and Overfitting in DenseNet with Specialized Dropout
- Optimization on Product Submanifolds of Convolution Kernels
- Super Interaction Neural Network
- Multi-scale Convolution Aggregation and Stochastic Feature Reuse for DenseNets
- Batch Normalization and the impact of batch structure on the behavior of deep convolution networks
- Selective Output Smoothing Regularization: Regularize Neural Networks by Softening Output Distributions
- Itsy Bitsy SpiderNet: Fully Connected Residual Network for Fraud Detection
- SelectScale: Mining More Patterns from Images via Selective and Soft Dropout
- DropRegion Training of Inception Font Network for High-Performance Chinese Font Recognition
- Deep Competitive Pathway Networks
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification
- Sequenced-Replacement Sampling for Deep Learning
- Structure Learning of Deep Neural Networks with Q-Learning
- Stochastic Deep Compressive Sensing for the Reconstruction of Diffusion Tensor Cardiac MRI
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
- Unbounded Output Networks for Classification