Deep Learning Meets Sparse Regularization: A Signal Processing Perspective
arXiv:2301.09554 · doi:10.1109/MSP.2023.3286988
Abstract
Deep learning has been wildly successful in practice and most state-of-the-art machine learning methods are based on neural networks. Lacking, however, is a rigorous mathematical theory that adequately explains the amazing performance of deep neural networks. In this article, we present a relatively new mathematical framework that provides the beginning of a deeper understanding of deep learning. This framework precisely characterizes the functional properties of neural networks that are trained to fit to data. The key mathematical tools which support this framework include transform-domain sparse regularization, the Radon transform of computed tomography, and approximation theory, which are all techniques deeply rooted in signal processing. This framework explains the effect of weight decay regularization in neural network training, the use of skip connections and low-rank weight matrices in network architectures, the role of sparsity in neural networks, and explains why neural networks can perform well in high-dimensional problems.
References in corpus (6)
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Sharp Bounds on the Approximation Rates, Metric Entropy, and -widths of Shallow Neural Networks
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- Near-Minimax Optimal Estimation With Shallow ReLU Neural Networks
- Ridges, Neural Networks, and the Radon Transform
- Variation Spaces for Multi-Output Neural Networks: Insights on Multi-Task Learning and Network Compression
Cited by in corpus (5)
- Distributional Extension and Invertibility of the -Plane Transform and Its Dual
- Weighted variation spaces and approximation by shallow ReLU networks
- Smoothing the Edges: Smooth Optimization for Sparse Regularization using Hadamard Overparametrization
- Function-Space Optimality of Neural Architectures with Multivariate Nonlinearities
- An Overview of Low-Rank Structures in the Training and Adaptation of Large Models