Measuring the Intrinsic Dimension of Objective Landscapes
arXiv:1804.08838
Abstract
Many recently trained neural networks employ large numbers of parameters to achieve good performance. One may intuitively use the number of parameters required as a rough gauge of the difficulty of a problem. But how accurate are such notions? How many parameters are really needed? In this paper we attempt to answer this question by training networks not in their native parameter space, but instead in a smaller, randomly oriented subspace. We slowly increase the dimension of this subspace, note at which dimension solutions first appear, and define this to be the intrinsic dimension of the objective landscape. The approach is simple to implement, computationally tractable, and produces several suggestive conclusions. Many problems have smaller intrinsic dimensions than one might suspect, and the intrinsic dimension for a given dataset varies little across a family of models with vastly different sizes. This latter result has the profound implication that once a parameter space is large enough to solve a problem, extra parameters serve directly to increase the dimensionality of the solution manifold. Intrinsic dimension allows some quantitative comparison of problem difficulty across supervised, reinforcement, and other types of learning where we conclude, for example, that solving the inverted pendulum problem is 100 times easier than classifying digits from MNIST, and playing Atari Pong from pixels is about as hard as classifying CIFAR-10. In addition to providing new cartography of the objective landscapes wandered by parameterized models, the method is a simple technique for constructively obtaining an upper bound on the minimum description length of a solution. A byproduct of this construction is a simple approach for compressing networks, in some cases by more than 100 times.
Published in ICLR 2018
References in corpus (4)
Cited by in corpus (35)
- A deep-learning-based surrogate model for data assimilation in dynamic subsurface flow problems
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Multi-task Self-Supervised Learning for Human Activity Detection
- An Intriguing Failing of Convolutional Neural Networks and the CoordConv Solution
- An Empirical Model of Large-Batch Training
- Soft-Label Dataset Distillation and Text Dataset Distillation
- Insights on representational similarity in neural networks with canonical correlation
- Superposition of many models into one
- Joint Embedding of Words and Labels for Text Classification
- Neural network interpretation using descrambler groups
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- What does it mean to understand a neural network?
- Adversarial Reprogramming of Neural Networks
- A Neural Scaling Law from the Dimension of the Data Manifold
- Subspace Inference for Bayesian Deep Learning
- Emergent properties of the local geometry of neural loss landscapes
- Large Scale Structure of Neural Network Loss Landscapes
- Can Unconditional Language Models Recover Arbitrary Sentences?
- Learning Neural Network Subspaces
- Dimension Estimation Using Autoencoders
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Multi-task neural networks by learned contextual inputs
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- On Scalable and Efficient Computation of Large Scale Optimal Transport
- Improving Generalization by Controlling Label-Noise Information in Neural Network Weights
- How many degrees of freedom do we need to train deep networks: a loss landscape perspective
- Privately Learning Subspaces
- Revisiting minimum description length complexity in overparameterized models
- Kernel Dependence Network
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Policy Manifold Search for Improving Diversity-based Neuroevolution