MixConv: Mixed Depthwise Convolutional Kernels
arXiv:1907.09595
Abstract
Depthwise convolution is becoming increasingly popular in modern efficient ConvNets, but its kernel size is often overlooked. In this paper, we systematically study the impact of different kernel sizes, and observe that combining the benefits of multiple kernel sizes can lead to better accuracy and efficiency. Based on this observation, we propose a new mixed depthwise convolution (MixConv), which naturally mixes up multiple kernel sizes in a single convolution. As a simple drop-in replacement of vanilla depthwise convolution, our MixConv improves the accuracy and efficiency for existing MobileNets on both ImageNet classification and COCO object detection. To demonstrate the effectiveness of MixConv, we integrate it into AutoML search space and develop a new family of models, named as MixNets, which outperform previous mobile models including MobileNetV2 [20] (ImageNet top-1 accuracy +4.2%), ShuffleNetV2 [16] (+3.5%), MnasNet [26] (+1.3%), ProxylessNAS [2] (+2.2%), and FBNet [27] (+2.0%). In particular, our MixNet-L achieves a new state-of-the-art 78.9% ImageNet top-1 accuracy under typical mobile settings (<600M FLOPS). Code is at https://github.com/ tensorflow/tpu/tree/master/models/official/mnasnet/mixnet
BMVC 2019
References in corpus (2)
Cited by in corpus (49)
- EfficientNetV2: Smaller Models and Faster Training
- DeepViT: Towards Deeper Vision Transformer
- Coordinate Attention for Efficient Mobile Network Design
- Neural Architecture Transfer
- Quantum machine learning for image classification
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition
- A New Dataset, Poisson GAN and AquaNet for Underwater Object Grabbing
- PP-LCNet: A Lightweight CPU Convolutional Neural Network
- Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation
- Zero-Cost Proxies for Lightweight NAS
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- AtomNAS: Fine-Grained End-to-End Neural Architecture Search
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
- Noisy Differentiable Architecture Search
- HS-ResNet: Hierarchical-Split Block on Convolutional Neural Network
- Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search
- DARTS-: Robustly Stepping out of Performance Collapse Without Indicators
- Blockwisely Supervised Neural Architecture Search with Knowledge Distillation
- PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System
- Learned Threshold Pruning
- LSQ+: Improving low-bit quantization through learnable offsets and better initialization
- Hyper-Convolution Networks for Biomedical Image Segmentation
- MixPath: A Unified Approach for One-shot Neural Architecture Search
- LeanConvNets: Low-cost Yet Effective Convolutional Neural Networks
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- Rethinking Channel Dimensions for Efficient Model Design
- Rethinking Bottleneck Structure for Efficient Mobile Network Design
- EdgeCNN: Convolutional Neural Network Classification Model with small inputs for Edge Computing
- Lite-HRNet: A Lightweight High-Resolution Network
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- XSepConv: Extremely Separated Convolution
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- S3NAS: Fast NPU-aware Neural Architecture Search Methodology
- Evolving Search Space for Neural Architecture Search
- HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass Lens
- A Deep Attentive Convolutional Neural Network for Automatic Cortical Plate Segmentation in Fetal MRI
- BS-NAS: Broadening-and-Shrinking One-Shot NAS with Searchable Numbers of Channels
- HybridNetSeg: A Compact Hybrid Network for Retinal Vessel Segmentation
- HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers
- SCARLET-NAS: Bridging the Gap between Stability and Scalability in Weight-sharing Neural Architecture Search
- RoBIC: A benchmark suite for assessing classifiers robustness
- Efficient Human Pose Estimation with Depthwise Separable Convolution and Person Centroid Guided Joint Grouping
- GreedyNASv2: Greedier Search with a Greedy Path Filter
- MixTConv: Mixed Temporal Convolutional Kernels for Efficient Action Recogntion
- MixFaceNets: Extremely Efficient Face Recognition Networks
- Malleable 2.5D Convolution: Learning Receptive Fields along the Depth-axis for RGB-D Scene Parsing
- One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space Shrinking