MS-DARTS: Mean-Shift Based Differentiable Architecture Search
arXiv:2108.09996
Abstract
Differentiable Architecture Search (DARTS) is an effective continuous relaxation-based network architecture search (NAS) method with low search cost. It has attracted significant attentions in Auto-ML research and becomes one of the most useful paradigms in NAS. Although DARTS can produce superior efficiency over traditional NAS approaches with better control of complex parameters, oftentimes it suffers from stabilization issues in producing deteriorating architectures when discretizing the continuous architecture. We observed considerable loss of validity causing dramatic decline in performance at this final discretization step of DARTS. To address this issue, we propose a Mean-Shift based DARTS (MS-DARTS) to improve stability based on sampling and perturbation. Our approach can improve bot the stability and accuracy of DARTS, by smoothing the loss landscape and sampling architecture parameters within a suitable bandwidth. We investigate the convergence of our mean-shift approach, together with the effects of bandwidth selection that affects stability and accuracy. Evaluations performed on CIFAR-10, CIFAR-100, and ImageNet show that MS-DARTS archives higher performance over other state-of-the-art NAS methods with reduced search cost.
14pages
References in corpus (12)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Neural Architecture Search with Reinforcement Learning
- Improved Regularization of Convolutional Neural Networks with Cutout
- Neural Architecture Search: A Survey
- DARTS: Differentiable Architecture Search
- Neural Architecture Optimization
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- Understanding and Robustifying Differentiable Architecture Search
- DARTS+: Improved Differentiable Architecture Search with Early Stopping
- NAS-Bench-1Shot1: Benchmarking and Dissecting One-shot Neural Architecture Search
- MaskConnect: Connectivity Learning by Gradient Descent
- Latency-Aware Differentiable Neural Architecture Search