Risk Bounds for High-dimensional Ridge Function Combinations Including Neural Networks
arXiv:1607.01434
Abstract
Let be a function on with an assumption of a spectral norm . For various noise settings, we show that , where is the sample size and is either a penalized least squares estimator or a greedily obtained version of such using linear combinations of sinusoidal, sigmoidal, ramp, ramp-squared or other smooth ridge functions. The candidate fits may be chosen from a continuum of functions, thus avoiding the rigidity of discretizations of the parameter space. On the other hand, if the candidate fits are chosen from a discretization, we show that . This work bridges non-linear and non-parametric function estimation and includes single-hidden layer nets. Unlike past theory for such settings, our bound shows that the risk is small even when the input dimension of an infinite-dimensional parameterized dictionary is much larger than the available sample size. When the dimension is larger than the cube root of the sample size, this quantity is seen to improve the more familiar risk bound of , also investigated here.
Submitted to Annals of Statistics
References in corpus (5)
- Approximation and learning by greedy algorithms
- Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
- Recovery Guarantees for One-hidden-layer Neural Networks
- Learning Halfspaces and Neural Networks with Random Initialization
- Learning Combinations of Sigmoids Through Gradient Estimation
Cited by in corpus (21)
- Nonparametric regression using deep neural networks with ReLU activation function
- A Theoretical Analysis of Deep Q-Learning
- A Priori Estimates of the Population Risk for Two-layer Neural Networks
- A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics
- Neural Policy Gradient Methods: Global Optimality and Rates of Convergence
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- A Priori Estimates of the Population Risk for Residual Networks
- Approximation and Estimation for High-Dimensional Deep Learning Networks
- A Selective Overview of Deep Learning
- Banach Space Representer Theorems for Neural Networks and Ridge Splines
- Representation formulas and pointwise properties for Barron functions
- On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics
- Complexity, Statistical Risk, and Metric Entropy of Deep Nets Using Total Path Variation
- Quantile regression with deep ReLU Networks: Estimators and minimax rates
- The Barron Space and the Flow-induced Function Spaces for Neural Network Models
- Deep Learning Partial Least Squares
- Spectral Pruning: Compressing Deep Neural Networks via Spectral Analysis and its Generalization Error
- A reduced order Schwarz method for nonlinear multiscale elliptic equations based on two-layer neural networks
- Embedding Inequalities for Barron-type Spaces
- A priori generalization error for two-layer ReLU neural network through minimum norm solution
- Minimax Lower Bounds for Ridge Combinations Including Neural Nets