Approximation spaces of deep neural networks
arXiv:1905.01208
Abstract
We study the expressivity of deep neural networks. Measuring a network's complexity by its number of connections or by its number of neurons, we consider the class of functions for which the error of best approximation with networks of a given complexity decays at a certain rate when increasing the complexity budget. Using results from classical approximation theory, we show that this class can be endowed with a (quasi)-norm that makes it a linear function space, called approximation space. We establish that allowing the networks to have certain types of "skip connections" does not change the resulting approximation spaces. We also discuss the role of the network's nonlinearity (also known as activation function) on the resulting spaces, as well as the role of depth. For the popular ReLU nonlinearity and its powers, we relate the newly constructed spaces to classical Besov spaces. The established embeddings highlight that some functions of very low Besov smoothness can nevertheless be well approximated by neural networks, if these networks are sufficiently deep.
Cited by in corpus (11)
- Deep Network Approximation for Smooth Functions
- Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth
- Full error analysis for the training of deep neural networks
- Learning with tree tensor networks: complexity estimates and model selection
- How degenerate is the parametrization of neural networks with the ReLU activation function?
- Approximation of Smoothness Classes by Deep Rectifier Networks
- NEU: A Meta-Algorithm for Universal UAP-Invariant Feature Representation
- Kolmogorov Width Decay and Poor Approximators in Machine Learning: Shallow Neural Networks, Random Feature Models and Neural Tangent Kernels
- Variable Selection with Rigorous Uncertainty Quantification using Deep Bayesian Neural Networks: Posterior Concentration and Bernstein-von Mises Phenomenon
- Translating Diffusion, Wavelets, and Regularisation into Residual Networks
- Error bounds for PDE-regularized learning