Collapse of Deep and Narrow Neural Nets
arXiv:1808.04947 · doi:10.4208/cicp.OA-2020-0165
Abstract
Recent theoretical work has demonstrated that deep neural networks have superior performance over shallow networks, but their training is more difficult, e.g., they suffer from the vanishing gradient problem. This problem can be typically resolved by the rectified linear unit (ReLU) activation. However, here we show that even for such activation, deep and narrow neural networks (NNs) will converge to erroneous mean or median states of the target function depending on the loss with high probability. Deep and narrow NNs are encountered in solving partial differential equations with high-order derivatives. We demonstrate this collapse of such NNs both numerically and theoretically, and provide estimates of the probability of collapse. We also construct a diagram of a safe region for designing NNs that avoid the collapse to erroneous states. Finally, we examine different ways of initialization and normalization that may avoid the collapse problem. Asymmetric initializations may reduce the probability of collapse but do not totally eliminate it.
Cited by in corpus (23)
- DeepONet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators
- Unsupervised Deep Learning for Massive MIMO Hybrid Beamforming
- ERANNs: Efficient Residual Audio Neural Networks for Audio Pattern Recognition
- Learning Specialized Activation Functions for Physics-informed Neural Networks
- A comprehensive and fair comparison of two neural operators (with practical extensions) based on FAIR data
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- Combining data assimilation and machine learning to estimate parameters of a convective-scale model
- CACTO: Continuous Actor-Critic with Trajectory Optimization -- Towards global optimality
- All-optical nonlinear activation function based on stimulated Brillouin scattering
- Modelling the galaxy-halo connection with semi-recurrent neural networks
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Quantum DeepONet: Neural operators accelerated by quantum computing
- Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models
- Polynomial Ridge Flowfield Estimation
- Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
- Dynamic Neural Diversification: Path to Computationally Sustainable Neural Networks
- A discrete physics-informed training for projection-based reduced order models with neural networks
- REAct: Rational Exponential Activation for Better Learning and Generalization in PINNs
- EffCNet: An Efficient CondenseNet for Image Classification on NXP BlueBox
- Bayesian optimization approach for tracking the location and orientation of a moving target using far-field data
- A deep learning approach to the texture optimization problem for friction control in lubricated contacts
- Leaky ReLUs That Differ in Forward and Backward Pass Facilitate Activation Maximization in Deep Neural Networks