Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov space
arXiv:1910.12799
Abstract
Deep learning has exhibited superior performance for various tasks, especially for high-dimensional datasets, such as images. To understand this property, we investigate the approximation and estimation ability of deep learning on anisotropic Besov spaces. The anisotropic Besov space is characterized by direction-dependent smoothness and includes several function classes that have been investigated thus far. We demonstrate that the approximation error and estimation error of deep learning only depend on the average value of the smoothness parameters in all directions. Consequently, the curse of dimensionality can be avoided if the smoothness of the target function is highly anisotropic. Unlike existing studies, our analysis does not require a low-dimensional structure of the input data. We also investigate the minimax optimality of deep learning and compare its performance with that of the kernel method (more generally, linear estimators). The results show that deep learning has better dependence on the input dimensionality if the target function possesses anisotropic smoothness, and it achieves an adaptive rate for functions with spatially inhomogeneous smoothness.
Accepted in NeurIPS2021
References in corpus (7)
- Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality
- Adaptive Approximation and Generalization of Deep Neural Network with Intrinsic Dimensionality
- Deep ReLU network approximation of functions on a manifold
- Rapid Convergence of the Unadjusted Langevin Algorithm: Isoperimetry Suffices
- On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
- Optimal Learning with Anisotropic Gaussian SVMs
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Cited by in corpus (5)
- Black-Box Optimization with Local Generative Surrogates
- Statistical theory for image classification using deep convolutional neural networks with cross-entropy loss under the hierarchical max-pooling model
- Particle Dual Averaging: Optimization of Mean Field Neural Networks with Global Convergence Rate Analysis
- Deep Regression for Repeated Measurements
- Estimation error analysis of deep learning on the regression problem on the variable exponent Besov space