A Priori Estimates of the Population Risk for Residual Networks
arXiv:1903.02154
Abstract
Optimal a priori estimates are derived for the population risk, also known as the generalization error, of a regularized residual network model. An important part of the regularized model is the usage of a new path norm, called the weighted path norm, as the regularization term. The weighted path norm treats the skip connections and the nonlinearities differently so that paths with more nonlinearities are regularized by larger weights. The error estimates are a priori in the sense that the estimates depend only on the target function, not on the parameters obtained in the training process. The estimates are optimal, in a high dimensional setting, in the sense that both the bound for the approximation and estimation errors are comparable to the Monte Carlo error rates. A crucial step in the proof is to establish an optimal bound for the Rademacher complexity of the residual networks. Comparisons are made with existing norm-based generalization error bounds.
References in corpus (1)
Cited by in corpus (7)
- A Selective Overview of Deep Learning
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
- Complexity Measures for Neural Networks with General Activation Functions Using Path-based Norms
- The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models
- Kolmogorov Width Decay and Poor Approximators in Machine Learning: Shallow Neural Networks, Random Feature Models and Neural Tangent Kernels