Representing smooth functions as compositions of near-identity functions with implications for deep network optimization
arXiv:1804.05012
Abstract
We show that any smooth bi-Lipschitz can be represented exactly as a composition of functions that are close to the identity in the sense that each is Lipschitz, and the Lipschitz constant decreases inversely with the number of functions composed. This implies that can be represented to any accuracy by a deep residual network whose nonlinear layers compute functions with a small Lipschitz constant. Next, we consider nonlinear regression with a composition of near-identity nonlinear maps. We show that, regarding Fréchet derivatives with respect to the , any critical point of a quadratic criterion in this near-identity region must be a global minimizer. In contrast, if we consider derivatives with respect to parameters of a fixed-size residual network with sigmoid activation functions, we show that there are near-identity critical points that are suboptimal, even in the realizable case. Informally, this means that functional gradient methods for residual networks cannot get stuck at suboptimal critical points corresponding to near-identity layers, whereas parametric gradient methods for sigmoidal residual networks suffer from suboptimal critical points in the near-identity region.
References in corpus (5)
Cited by in corpus (11)
- Data Driven Governing Equations Approximation Using Deep Neural Networks
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- GResNet: Graph Residual Network for Reviving Deep GNNs from Suspended Animation
- A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth
- A case for new neural network smoothness constraints
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Are deep ResNets provably better than linear predictors?
- DeDUCE: Generating Counterfactual Explanations Efficiently
- Chaining Meets Chain Rule: Multilevel Entropic Regularization and Training of Neural Nets
- A class of robust numerical methods for solving dynamical systems with multiple time scales
- Step Size Matters in Deep Learning