Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
arXiv:2011.06702 · doi:10.1109/BigData50022.2020.9378359
Abstract
Modern deep neural network (DNN) trainings utilize various training techniques, e.g., nonlinear activation functions, batch normalization, skip-connections, etc. Despite their effectiveness, it is still mysterious how they help accelerate DNN trainings in practice. In this paper, we provide an empirical study of the regularization effect of these training techniques on DNN optimization. Specifically, we find that the optimization trajectories of successful DNN trainings consistently obey a certain regularity principle that regularizes the model update direction to be aligned with the trajectory direction. Theoretically, we show that such a regularity principle leads to a convergence guarantee in nonconvex optimization and the convergence rate depends on a regularization parameter. Empirically, we find that DNN trainings that apply the training techniques achieve a fast convergence and obey the regularity principle with a large regularization parameter, implying that the model updates are well aligned with the trajectory. On the other hand, DNN trainings without the training techniques have slow convergence and obey the regularity principle with a small regularization parameter, implying that the model updates are not well aligned with the trajectory. Therefore, different training techniques regularize the model update direction via the regularity principle to facilitate the convergence.
9 pages, 16 figures, this paper has been accepted as a short paper by the conference of IEEE-bigdata-2020
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- On the Convergence of Adam and Beyond
- Convergence Analysis of Two-layer Neural Networks with ReLU Activation
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Gradient Descent Finds Global Minima of Deep Neural Networks
- Recovery Guarantees for One-hidden-layer Neural Networks
- A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks
- On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
- Skip Connections Eliminate Singularities
- Learning One-hidden-layer ReLU Networks via Gradient Descent
- Characterization of Gradient Dominance and Regularity Conditions for Neural Networks
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path