Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality
arXiv:1810.08033
Abstract
Deep learning has shown high performances in various types of tasks from visual recognition to natural language processing, which indicates superior flexibility and adaptivity of deep learning. To understand this phenomenon theoretically, we develop a new approximation and estimation error analysis of deep learning with the ReLU activation for functions in a Besov space and its variant with mixed smoothness. The Besov space is a considerably general function space including the Holder space and Sobolev space, and especially can capture spatial inhomogeneity of smoothness. Through the analysis in the Besov space, it is shown that deep learning can achieve the minimax optimal rate and outperform any non-adaptive (linear) estimator such as kernel ridge regression, which shows that deep learning has higher adaptivity to the spatial inhomogeneity of the target function than other estimators such as linear ones. In addition to this, it is shown that deep learning can avoid the curse of dimensionality if the target function is in a mixed smooth Besov space. We also show that the dependency of the convergence rate on the dimensionality is tight due to its minimax optimality. These results support high adaptivity of deep learning and its superior ability as a feature extractor.
Cited by in corpus (39)
- Deep Network Approximation for Smooth Functions
- A Theoretical Analysis of Deep Q-Learning
- Nonlinear Approximation via Compositions
- SelectNet: Self-paced Learning for High-dimensional Partial Differential Equations
- Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth
- Convergence Rate Analysis for Deep Ritz Method
- Deep ReLU network approximation of functions on a manifold
- Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators
- Near-Minimax Optimal Estimation With Shallow ReLU Neural Networks
- On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
- A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov space
- Exponential ReLU Neural Network Approximation Rates for Point and Edge Singularities
- Learning with tree tensor networks: complexity estimates and model selection
- Approximation in shift-invariant spaces with deep ReLU neural networks
- Convergence Rates of Variational Inference in Sparse Deep Learning
- Deep Nonparametric Regression on Approximate Manifolds: Non-Asymptotic Error Bounds with Polynomial Prefactors
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network
- Approximation of Smoothness Classes by Deep Rectifier Networks
- Optimal Nonparametric Inference via Deep Neural Network
- Quantile regression with deep ReLU Networks: Estimators and minimax rates
- On generalization bounds for deep networks based on loss surface implicit regularization
- RoeNets: Predicting Discontinuity of Hyperbolic Systems from Continuous Data
- On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning
- Misclassification bounds for PAC-Bayesian sparse deep learning
- Nonconvex sparse regularization for deep neural networks and its optimality
- 50 years since the Marr, Ito, and Albus models of the cerebellum
- Particle Dual Averaging: Optimization of Mean Field Neural Networks with Global Convergence Rate Analysis
- Approximation with Neural Networks in Variable Lebesgue Spaces
- Robust Density Estimation under Besov IPM Losses
- Sample Complexity of Offline Reinforcement Learning with Deep ReLU Networks
- Collocation approximation by deep neural ReLU networks for parametric elliptic PDEs with lognormal inputs
- Fast generalization error bound of deep learning without scale invariance of activation functions
- High-Dimensional Non-Parametric Density Estimation in Mixed Smooth Sobolev Spaces
- Theory of Deep Convolutional Neural Networks III: Approximating Radial Functions
- Deep Regression for Repeated Measurements
- Neural Estimation of Statistical Divergences
- Estimation error analysis of deep learning on the regression problem on the variable exponent Besov space
- Theory of Deep Convolutional Neural Networks II: Spherical Analysis