Generalization Guarantees for Neural Architecture Search with Train-Validation Split
arXiv:2104.14132
Abstract
Neural Architecture Search (NAS) is a popular method for automatically designing optimized architectures for high-performance deep learning. In this approach, it is common to use bilevel optimization where one optimizes the model weights over the training data (inner problem) and various hyperparameters such as the configuration of the architecture over the validation data (outer problem). This paper explores the statistical aspects of such problems with train-validation splits. In practice, the inner problem is often overparameterized and can easily achieve zero loss. Thus, a-priori it seems impossible to distinguish the right hyperparameters based on training loss alone which motivates a better understanding of the role of train-validation split. To this aim this work establishes the following results. (1) We show that refined properties of the validation loss such as risk and hyper-gradients are indicative of those of the true test loss. This reveals that the outer problem helps select the most generalizable model and prevent overfitting with a near-minimal validation sample size. This is established for continuous search spaces which are relevant for differentiable schemes. Extensions to transfer learning are developed in terms of the mismatch between training & validation distributions. (2) We establish generalization bounds for NAS problems with an emphasis on an activation search problem. When optimized with gradient-descent, we show that the train-validation procedure returns the best (model, architecture) pair even if all architectures can perfectly fit the training data to achieve zero error. (3) Finally, we highlight connections between NAS, multiple kernel learning, and low-rank matrix learning. The latter leads to novel insights where the solution of the outer problem can be accurately learned via efficient spectral methods to achieve near-minimal risk.
ICML 2021
References in corpus (25)
- Neural Architecture Search with Reinforcement Learning
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- On Exact Computation with an Infinitely Wide Neural Net
- Exploring Generalization in Deep Learning
- SNAS: Stochastic Neural Architecture Search
- On Lazy Training in Differentiable Programming
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics
- Enhanced Convolutional Neural Tangent Kernels
- Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Neural Kernels Without Tangents
- AutoHAS: Efficient Hyperparameter and Architecture Search
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems
- Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems
- Towards NNGP-guided Neural Architecture Search
- How Important is the Train-Validation Split in Meta-Learning?
- Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
- Geometry-Aware Gradient Algorithms for Neural Architecture Search
- Rademacher upper bounds for cross-validation errors with an application to the lasso