Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations
arXiv:1710.10121
Abstract
In our work, we bridge deep neural network design with numerical differential equations. We show that many effective networks, such as ResNet, PolyNet, FractalNet and RevNet, can be interpreted as different numerical discretizations of differential equations. This finding brings us a brand new perspective on the design of effective deep architectures. We can take advantage of the rich knowledge in numerical analysis to guide us in designing new and potentially more effective deep networks. As an example, we propose a linear multi-step architecture (LM-architecture) which is inspired by the linear multi-step method solving ordinary differential equations. The LM-architecture is an effective structure that can be used on any ResNet-like networks. In particular, we demonstrate that LM-ResNet and LM-ResNeXt (i.e. the networks obtained by applying the LM-architecture on ResNet and ResNeXt respectively) can achieve noticeably higher accuracy than ResNet and ResNeXt on both CIFAR and ImageNet with comparable numbers of trainable parameters. In particular, on both CIFAR and ImageNet, LM-ResNet/LM-ResNeXt can significantly compress (\%) the original networks while maintaining a similar performance. This can be explained mathematically using the concept of modified equation from numerical analysis. Last but not least, we also establish a connection between stochastic control and noise injection in the training process which helps to improve generalization of the networks. Furthermore, by relating stochastic training strategy with stochastic dynamic system, we can easily apply stochastic training to the networks with the LM-architecture. As an example, we introduced stochastic depth to LM-ResNet and achieve significant improvement over the original LM-ResNet on CIFAR10.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Shake-Shake regularization
- PDE-Net: Learning PDEs from Data
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Highway and Residual Networks learn Unrolled Iterative Estimation
- DiracNets: Training Very Deep Neural Networks Without Skip-Connections
- Reversible Architectures for Arbitrarily Deep Residual Neural Networks
- Deep Residual Learning and PDEs on Manifold
- Information Dropout: Learning Optimal Representations Through Noisy Computation
Cited by in corpus (34)
- Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network
- PDE-Net 2.0: Learning PDEs from Data with A Numeric-Symbolic Hybrid Deep Network
- A deep learning enabler for non-intrusive reduced order modeling of fluid flows
- An artificial neural network framework for reduced order modeling of transient flows
- Multistep Neural Networks for Data-driven Discovery of Nonlinear Dynamical Systems
- MgNet: A Unified Framework of Multigrid and Convolutional Neural Network
- TeaNet: universal neural network interatomic potential inspired by iterative electronic relaxations
- STDEN: Towards Physics-Guided Neural Networks for Traffic Flow Prediction
- FEA-Net: A Physics-guided Data-driven Model for Efficient Mechanical Response Prediction
- PDE-based Group Equivariant Convolutional Neural Networks
- Combining distribution-based neural networks to predict weather forecast probabilities
- Physics-Informed Supervised Residual Learning for Electromagnetic Modeling
- Machine Learning from a Continuous Viewpoint
- A Helmholtz equation solver using unsupervised learning: Application to transcranial ultrasound
- Evolutionary Preference Learning via Graph Nested GRU ODE for Session-based Recommendation
- Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE
- Galaxy Morphology Classification using Neural Ordinary Differential Equations
- A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth
- Recurrent Neural Networks in the Eye of Differential Equations
- NeuPDE: Neural Network Based Ordinary and Partial Differential Equations for Modeling Time-Dependent Data
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
- LeanConvNets: Low-cost Yet Effective Convolutional Neural Networks
- State Space Representations of Deep Neural Networks
- The gap between theory and practice in function approximation with deep neural networks
- Layer-Parallel Training of Deep Residual Neural Networks
- Feature Flow Regularization: Improving Structured Sparsity in Deep Neural Networks
- Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
- Neural Dynamics on Complex Networks
- Translating Diffusion, Wavelets, and Regularisation into Residual Networks
- Neural ODEs with stochastic vector field mixtures
- Time-Continuous Energy-Conservation Neural Network for Structural Dynamics Analysis
- A Differential Game Theoretic Neural Optimizer for Training Residual Networks
- A Grid-Structured Model of Tubular Reactors
- Machine Learning Methods for Autonomous Ordinary Differential Equations