Approximation by Combinations of ReLU and Squared ReLU Ridge Functions with and Controls
arXiv:1607.07819
Abstract
We establish and error bounds for functions of many variables that are approximated by linear combinations of ReLU (rectified linear unit) and squared ReLU ridge functions with and controls on their inner and outer parameters. With the squared ReLU ridge function, we show that the approximation error is inversely proportional to the inner layer sparsity and it need only be sublinear in the outer layer sparsity. Our constructions are obtained using a variant of the Jones-Barron probabilistic method, which can be interpreted as either stratified sampling with proportionate allocation or two-stage cluster sampling. We also provide companion error lower bounds that reveal near optimality of our constructions. Despite the sparsity assumptions, we showcase the richness and flexibility of these ridge combinations by defining a large family of functions, in terms of certain spectral conditions, that are particularly well approximated by them.
References in corpus (1)
Cited by in corpus (7)
- Banach Space Representer Theorems for Neural Networks and Ridge Splines
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Universality of Deep Convolutional Neural Networks
- Approximation Rates for Neural Networks with General Activation Functions
- Spectral Pruning: Compressing Deep Neural Networks via Spectral Analysis and its Generalization Error
- Super-resolution meets machine learning: approximation of measures
- Function approximation with zonal function networks with activation functions analogous to the rectified linear unit functions