What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
arXiv:2105.03361 · doi:10.1137/21M1418642
Abstract
We develop a variational framework to understand the properties of functions learned by fitting deep neural networks with rectified linear unit activations to data. We propose a new function space, which is reminiscent of classical bounded variation-type spaces, that captures the compositional structure associated with deep neural networks. We derive a representer theorem showing that deep ReLU networks are solutions to regularized data fitting problems over functions from this space. The function space consists of compositions of functions from the Banach spaces of second-order bounded variation in the Radon domain. These are Banach spaces with sparsity-promoting norms, giving insight into the role of sparsity in deep neural networks. The neural network solutions have skip connections and rank bounded weight matrices, providing new theoretical support for these common architectural choices. The variational problem we study can be recast as a finite-dimensional neural network training problem with regularization schemes related to the notions of weight decay and path-norm regularization. Finally, our analysis builds on techniques from variational spline theory, providing new connections between deep neural networks and splines.
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Norm-Based Capacity Control in Neural Networks
- Complexity, Statistical Risk, and Metric Entropy of Deep Nets Using Total Path Variation
Cited by in corpus (10)
- Near-Minimax Optimal Estimation With Shallow ReLU Neural Networks
- Deep Learning Meets Sparse Regularization: A Signal Processing Perspective
- Optimal rates of approximation by shallow ReLU neural networks and applications to nonparametric regression
- Spectral Barron space for deep neural network approximation
- Weighted variation spaces and approximation by shallow ReLU networks
- Distributional Extension and Invertibility of the -Plane Transform and Its Dual
- On extremal points for some vectorial total variation seminorms
- Function-Space Optimality of Neural Architectures with Multivariate Nonlinearities
- The Directional Bias Helps Stochastic Gradient Descent to Generalize in Kernel Regression Models
- Ridgeless Interpolation with Shallow ReLU Networks in is Nearest Neighbor Curvature Extrapolation and Provably Generalizes on Lipschitz Functions