Neural Network with Unbounded Activation Functions is Universal Approximator
arXiv:1505.03654 · doi:10.1016/j.acha.2015.12.005
Abstract
This paper presents an investigation of the approximation property of neural networks with unbounded activation functions, such as the rectified linear unit (ReLU), which is the new de-facto standard of deep learning. The ReLU network can be analyzed by the ridgelet transform with respect to Lizorkin distributions. By showing three reconstruction formulas by using the Fourier slice theorem, the Radon transform, and Parseval's relation, it is shown that a neural network with unbounded activation functions still satisfies the universal approximation property. As an additional consequence, the ridgelet transform, or the backprojection filter in the Radon domain, is what the network learns after backpropagation. Subject to a constructive admissibility condition, the trained network can be obtained by simply discretizing the ridgelet transform, without backpropagation. Numerical examples not only support the consistency of the admissibility condition but also imply that some non-admissible cases result in low-pass filtering.
under review; first revised version
References in corpus (1)
Cited by in corpus (28)
- An overview of deep learning in medical imaging focusing on MRI
- A survey on modern trainable activation functions
- Geometric deep learning for computational mechanics Part I: Anisotropic Hyperelasticity
- ReLU Networks as Surrogate Models in Mixed-Integer Linear Programs
- Machine Learning from a Continuous Viewpoint
- A machine learning aided global diagnostic and comparative tool to assess effect of quarantine control in Covid-19 spread
- A simple and efficient architecture for trainable activation functions
- Regression methods in waveform modeling: a comparative study
- On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
- Physical deep learning based on optimal control of dynamical systems
- Deep learning fluid flow reconstruction around arbitrary two-dimensional objects from sparse sensors using conformal mappings
- Neural Fractional Differential Equations
- Entanglement Features of Random Neural Network Quantum States
- BAMCAFE: A Bayesian Machine Learning Advanced Forecast Ensemble Method for Complex Turbulent Systems with Partial Observations
- Autoencoder-driven Spiral Representation Learning for Gravitational Wave Surrogate Modelling
- Next2You: Robust Copresence Detection Based on Channel State Information
- Symmetry & critical points for a model shallow neural network
- Growing Cosine Unit: A Novel Oscillatory Activation Function That Can Speedup Training and Reduce Parameters in Convolutional Neural Networks
- A unified Fourier slice method to derive ridgelet transform for a variety of depth-2 neural networks
- Distributional Extension and Invertibility of the -Plane Transform and Its Dual
- Fredholm Neural Networks
- Disentangling private classes through regularization
- Data-driven Initial Gap Identification of Piecewise-linear Systems using Sparse Regression and Universal Approximation Theorem
- Aspects of importance sampling in parameter selection for neural networks using ridgelet transform
- Interpreting Deep Neural Network-Based Receiver Under Varying Signal-To-Noise Ratios
- Bayesian neural networks with interpretable priors from Mercer kernels
- Piecewise Linear Functions Representable with Infinite Width Shallow ReLU Neural Networks
- Rapid Risk Minimization with Bayesian Models Through Deep Learning Approximation