In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
arXiv:1412.6614
Abstract
We present experiments demonstrating that some other form of capacity control, different from network size, plays a central role in learning multilayer feed-forward networks. We argue, partially through analogy to matrix factorization, that this is an inductive bias that can help shed light on deep learning.
9 pages, 2 figures
Cited by in corpus (27)
- Exploring Generalization in Deep Learning
- The Modern Mathematics of Deep Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Understanding the Failure Modes of Out-of-Distribution Generalization
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- The worst of both worlds: A comparative analysis of errors in learning from data in psychology and machine learning
- Deep Learning Meets Sparse Regularization: A Signal Processing Perspective
- Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
- SINR: Deconvolving Circular SAS Images Using Implicit Neural Representations
- Distance-Based Learning from Errors for Confidence Calibration
- The Low-Rank Simplicity Bias in Deep Networks
- Probing the Purview of Neural Networks via Gradient Analysis
- Should attention be all we need? The epistemic and ethical implications of unification in machine learning
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Regression as Classification: Influence of Task Formulation on Neural Network Features
- A Learning Theoretic Perspective on Local Explainability
- Pre-interpolation loss behaviour in neural networks
- On the approximation of functions by tanh neural networks
- Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape
- Rethinking Gauss-Newton for learning over-parameterized models
- Informative regularization for a multi-layer perceptron RR Lyrae classifier under data shift
- Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
- Implicit regularization of deep residual networks towards neural ODEs
- Painless step size adaptation for SGD
- On the Regularization of Autoencoders
- A Functional Perspective on Learning Symmetric Functions with Neural Networks
- Inherent Noise in Gradient Based Methods