On the existence of optimal shallow feedforward networks with ReLU activation
arXiv:2303.03950 · doi:10.4208/jml.230903
Abstract
We prove existence of global minima in the loss landscape for the approximation of continuous target functions using shallow feedforward artificial neural networks with ReLU activation. This property is one of the fundamental artifacts separating ReLU from other commonly used activation functions. We propose a kind of closure of the search space so that in the extended space minimizers exist. In a second step, we show under mild assumptions that the newly added functions in the extension perform worse than appropriate representable ReLU networks. This then implies that the optimal response in the extended target space is indeed the response of a ReLU network.
arXiv admin note: substantial text overlap with arXiv:2302.14690
References in corpus (4)
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex Problems
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- Cooling down stochastic differential equations: almost sure convergence
- Blow up phenomena for gradient descent optimization methods in the training of artificial neural networks