Zero-Cost Operation Scoring in Differentiable Architecture Search
arXiv:2106.06799
Abstract
We formalize and analyze a fundamental component of differentiable neural architecture search (NAS): local "operation scoring" at each operation choice. We view existing operation scoring functions as inexact proxies for accuracy, and we find that they perform poorly when analyzed empirically on NAS benchmarks. From this perspective, we introduce a novel \textit{perturbation-based zero-cost operation scoring} (Zero-Cost-PT) approach, which utilizes zero-cost proxies that were recently studied in multi-trial NAS but degrade significantly on larger search spaces, typical for differentiable NAS. We conduct a thorough empirical evaluation on a number of NAS benchmarks and large search spaces, from NAS-Bench-201, NAS-Bench-1Shot1, NAS-Bench-Macro, to DARTS-like and MobileNet-like spaces, showing significant improvements in both search time and accuracy. On the ImageNet classification task on the DARTS search space, our approach improved accuracy compared to the best current training-free methods (TE-NAS) while being over 10 faster (total searching time 25 minutes on a single GPU), and observed significantly better transferability on architectures searched on the CIFAR-10 dataset with an accuracy increase of 1.8 pp. Our code is available at: https://github.com/zerocostptnas/zerocost_operation_score.
Accepted at AAAI 2023
References in corpus (15)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Neural Architecture Search with Reinforcement Learning
- DARTS: Differentiable Architecture Search
- EfficientNetV2: Smaller Models and Faster Training
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Neural Architecture Optimization
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- Understanding and Robustifying Differentiable Architecture Search
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Faster gaze prediction with dense networks and Fisher pruning
- BRP-NAS: Prediction-based NAS using GCNs
- Rethinking Architecture Selection in Differentiable NAS
- Neural Predictor for Neural Architecture Search
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- Stronger NAS with Weaker Predictors