AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
arXiv:1903.11728
Abstract
We study how to set channel numbers in a neural network to achieve better accuracy under constrained resources (e.g., FLOPs, latency, memory footprint or model size). A simple and one-shot solution, named AutoSlim, is presented. Instead of training many network samples and searching with reinforcement learning, we train a single slimmable network to approximate the network accuracy of different channel configurations. We then iteratively evaluate the trained slimmable model and greedily slim the layer with minimal accuracy drop. By this single pass, we can obtain the optimized channel configurations under different resource constraints. We present experiments with MobileNet v1, MobileNet v2, ResNet-50 and RL-searched MNasNet on ImageNet classification. We show significant improvements over their default channel configurations. We also achieve better accuracy than recent channel pruning methods and neural architecture search methods. Notably, by setting optimized channel numbers, our AutoSlim-MobileNet-v2 at 305M FLOPs achieves 74.2% top-1 accuracy, 2.4% better than default MobileNet-v2 (301M FLOPs), and even 0.2% better than RL-searched MNasNet (317M FLOPs). Our AutoSlim-ResNet-50 at 570M FLOPs, without depthwise convolutions, achieves 1.3% better accuracy than MobileNet-v1 (569M FLOPs). Code and models will be available at: https://github.com/JiahuiYu/slimmable_networks
tech report
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Architecture Search with Reinforcement Learning
- Pruning Filters for Efficient ConvNets
- Emergence of Locomotion Behaviours in Rich Environments
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Slimmable Neural Networks
Cited by in corpus (35)
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Optimization for deep learning: theory and algorithms
- HRel: Filter Pruning based on High Relevance between Activation Maps and Class Labels
- Contrastive Self-supervised Neural Architecture Search
- Searching for Low-Bit Weights in Quantized Neural Networks
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Locally Free Weight Sharing for Network Width Search
- AdaSpring: Context-adaptive and Runtime-evolutionary Deep Model Compression for Mobile Applications
- Switchable Precision Neural Networks
- Pruning Self-attentions into Convolutional Layers in Single Path
- DMCP: Differentiable Markov Channel Pruning for Neural Networks
- Pruning via Iterative Ranking of Sensitivity Statistics
- EagleEye: Fast Sub-net Evaluation for Efficient Neural Network Pruning
- Searching for Efficient Multi-Stage Vision Transformers
- An Information Theory-inspired Strategy for Automatic Network Pruning
- Cascaded channel pruning using hierarchical self-distillation
- K-shot NAS: Learnable Weight-Sharing for NAS with K-shot Supernets
- Dynamic-OFA: Runtime DNN Architecture Switching for Performance Scaling on Heterogeneous Embedded Platforms
- Recursive-NeRF: An Efficient and Dynamically Growing NeRF
- HALP: Hardware-Aware Latency Pruning
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- NetAdaptV2: Efficient Neural Architecture Search with Fast Super-Network Training and Architecture Optimization
- Slimmable Generative Adversarial Networks
- Towards Efficient Convolutional Network Models with Filter Distribution Templates
- Joint Channel and Weight Pruning for Model Acceleration on Moblie Devices
- AdaPruner: Adaptive Channel Pruning and Effective Weights Inheritance
- BCNet: Searching for Network Width with Bilaterally Coupled Network
- Dynamic Slimmable Denoising Network
- Neural Inheritance Relation Guided One-Shot Layer Assignment Search
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- Width Transfer: On the (In)variance of Width Optimization
- Out-of-the-box channel pruned networks
- Masked Training of Neural Networks with Partial Gradients
- Prioritized Subnet Sampling for Resource-Adaptive Supernet Training