Exploring the Loss Landscape in Neural Architecture Search
arXiv:2005.02960
Abstract
Neural architecture search (NAS) has seen a steep rise in interest over the last few years. Many algorithms for NAS consist of searching through a space of architectures by iteratively choosing an architecture, evaluating its performance by training it, and using all prior evaluations to come up with the next choice. The evaluation step is noisy - the final accuracy varies based on the random initialization of the weights. Prior work has focused on devising new search algorithms to handle this noise, rather than quantifying or understanding the level of noise in architecture evaluations. In this work, we show that (1) the simplest hill-climbing algorithm is a powerful baseline for NAS, and (2), when the noise in popular NAS benchmark datasets is reduced to a minimum, hill-climbing to outperforms many popular state-of-the-art algorithms. We further back up this observation by showing that the number of local minima is substantially reduced as the noise decreases, and by giving a theoretical characterization of the performance of local search in NAS. Based on our findings, for NAS research we suggest (1) using local search as a baseline, and (2) denoising the training pipeline when possible.
References in corpus (9)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Neural Architecture Search with Reinforcement Learning
- NAS-Bench-101: Towards Reproducible Neural Architecture Search
- Random Search and Reproducibility for Neural Architecture Search
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- Best Practices for Scientific Research on Neural Architecture Search
- AlphaX: eXploring Neural Architectures with Deep Neural Networks and Monte Carlo Tree Search
- How Powerful are Performance Predictors in Neural Architecture Search?
- BANANAS: Bayesian Optimization with Neural Architectures for Neural Architecture Search