MobileNetV2: Inverted Residuals and Linear Bottlenecks
arXiv:1801.04381
Abstract
In this paper we describe a new mobile architecture, MobileNetV2, that improves the state of the art performance of mobile models on multiple tasks and benchmarks as well as across a spectrum of different model sizes. We also describe efficient ways of applying these mobile models to object detection in a novel framework we call SSDLite. Additionally, we demonstrate how to build mobile semantic segmentation models through a reduced form of DeepLabv3 which we call Mobile DeepLabv3. The MobileNetV2 architecture is based on an inverted residual structure where the input and output of the residual block are thin bottleneck layers opposite to traditional residual models which use expanded representations in the input an MobileNetV2 uses lightweight depthwise convolutions to filter features in the intermediate expansion layer. Additionally, we find that it is important to remove non-linearities in the narrow layers in order to maintain representational power. We demonstrate that this improves performance and provide an intuition that led to this design. Finally, our approach allows decoupling of the input/output domains from the expressiveness of the transformation, which provides a convenient framework for further analysis. We measure our performance on Imagenet classification, COCO object detection, VOC image segmentation. We evaluate the trade-offs between accuracy, and number of operations measured by multiply-adds (MAdd), as well as the number of parameters
References in corpus (6)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Neural Architecture Search with Reinforcement Learning
- Scalable Bayesian Optimization Using Deep Neural Networks
- The Power of Sparsity in Convolutional Neural Networks
- Design of Efficient Convolutional Layers using Single Intra-channel Convolution, Topological Subdivisioning and Spatial "Bottleneck" Structure
- Connectivity Learning in Multi-Branch Networks
Cited by in corpus (234)
- Knowledge Distillation: A Survey
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
- LocalViT: Analyzing Locality in Vision Transformers
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- MCUNet: Tiny Deep Learning on IoT Devices
- Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning
- Towards an Effective and Efficient Deep Learning Model for COVID-19 Patterns Detection in X-ray Images
- Less is More: Lighter and Faster Deep Neural Architecture for Tomato Leaf Disease Classification
- Evaluating the Single-Shot MultiBox Detector and YOLO Deep Learning Models for the Detection of Tomatoes in a Greenhouse
- Spatial Group-wise Enhance: Improving Semantic Feature Learning in Convolutional Networks
- Bringing AI To Edge: From Deep Learning's Perspective
- Insect pest image detection and recognition based on bio-inspired methods
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Adversarially Robust Distillation
- Network Pruning via Transformable Architecture Search
- Edge Intelligence: Architectures, Challenges, and Applications
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- Boosted EfficientNet: Detection of Lymph Node Metastases in Breast Cancer Using Convolutional Neural Network
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- Study of Different Deep Learning Approach with Explainable AI for Screening Patients with COVID-19 Symptoms: Using CT Scan and Chest X-ray Image Dataset
- Measuring the Algorithmic Efficiency of Neural Networks
- CGNet: A Light-weight Context Guided Network for Semantic Segmentation
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- OFFSEG: A Semantic Segmentation Framework For Off-Road Driving
- A Survey on Deep Learning for Polyp Segmentation: Techniques, Challenges and Future Trends
- Deep Level Sets: Implicit Surface Representations for 3D Shape Inference
- Unified Learning Approach for Egocentric Hand Gesture Recognition and Fingertip Detection
- SNIPER: Efficient Multi-Scale Training
- LegoDNN: Block-grained Scaling of Deep Neural Networks for Mobile Vision
- MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
- Oort: Efficient Federated Learning via Guided Participant Selection
- Detecting Visual Design Principles in Art and Architecture through Deep Convolutional Neural Networks
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- Refiner: Refining Self-attention for Vision Transformers
- WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection
- LogME: Practical Assessment of Pre-trained Models for Transfer Learning
- Efficiently utilizing complex-valued PolSAR image data via a multi-task deep learning framework
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Comparative evaluation of CNN architectures for Image Caption Generation
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- sharpDARTS: Faster and More Accurate Differentiable Architecture Search
- PFLD: A Practical Facial Landmark Detector
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- Pay Attention to MLPs
- Towards Better Accuracy-efficiency Trade-offs: Divide and Co-training
- Multiplicative noise and heavy tails in stochastic optimization
- Enabling Homomorphically Encrypted Inference for Large DNN Models
- DSConv: Efficient Convolution Operator
- EdgeSpeechNets: Highly Efficient Deep Neural Networks for Speech Recognition on the Edge
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack Transferability
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification
- FedBE: Making Bayesian Model Ensemble Applicable to Federated Learning
- RODEO: Replay for Online Object Detection
- Improving Object Detection from Scratch via Gated Feature Reuse
- Deep fusion of gray level co-occurrence matrices for lung nodule classification
- AutoHAS: Efficient Hyperparameter and Architecture Search
- AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT Systems
- Deep Learning for Identifying Iran's Cultural Heritage Buildings in Need of Conservation Using Image Classification and Grad-CAM
- Quantization for Rapid Deployment of Deep Neural Networks
- Optimized Three Deep Learning Models Based-PSO Hyperparameters for Beijing PM2.5 Prediction
- APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
- Aurora Guard: Real-Time Face Anti-Spoofing via Light Reflection
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Wisdom of Committees: An Overlooked Approach To Faster and More Accurate Models
- LOTR: Face Landmark Localization Using Localization Transformer
- UAVs Beneath the Surface: Cooperative Autonomy for Subterranean Search and Rescue in DARPA SubT
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
- Improved Techniques for Training Adaptive Deep Networks
- Marvel: A Data-centric Compiler for DNN Operators on Spatial Accelerators
- CovidExpert: A Triplet Siamese Neural Network framework for the detection of COVID-19
- Prospects for future studies using deep imaging: Analysis of individual Galactic cirrus filaments
- Rethinking Co-design of Neural Architectures and Hardware Accelerators
- Lightweight Regression Model with Prediction Interval Estimation for Computer Vision-based Winter Road Surface Condition Monitoring
- PowerNorm: Rethinking Batch Normalization in Transformers
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Frustratingly Easy Person Re-Identification: Generalizing Person Re-ID in Practice
- DNA: Differentiable Network-Accelerator Co-Search
- Understanding Reuse, Performance, and Hardware Cost of DNN Dataflows: A Data-Centric Approach Using MAESTRO
- Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection
- Anytime Inference with Distilled Hierarchical Neural Ensembles
- BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
- Music theme recognition using CNN and self-attention
- Fully Automatic Wound Segmentation with Deep Convolutional Neural Networks
- FOD-A: A Dataset for Foreign Object Debris in Airports
- Accelerating Sparse Deep Neural Networks
- Neural Transformers for Intraductal Papillary Mucosal Neoplasms (IPMN) Classification in MRI images
- Fine-Grained Neural Architecture Search
- AdversarialNAS: Adversarial Neural Architecture Search for GANs
- Learning to play the Chess Variant Crazyhouse above World Champion Level with Deep Neural Networks and Human Data
- See More Than Once -- Kernel-Sharing Atrous Convolution for Semantic Segmentation
- DSNet for Real-Time Driving Scene Semantic Segmentation
- The problem of dust attenuation in photometric decomposition of edge-on galaxies and possible solutions
- Deep Reasoning with Multi-Scale Context for Salient Object Detection
- Inspector Gadget: A Data Programming-based Labeling System for Industrial Images
- Low-latency Perception in Off-Road Dynamical Low Visibility Environments
- Distilling Knowledge via Knowledge Review
- Context-Aware Image Matting for Simultaneous Foreground and Alpha Estimation
- Neural Architecture Design for GPU-Efficient Networks
- Aurora Guard: Reliable Face Anti-Spoofing via Mobile Lighting System
- Real-Time High-Resolution Background Matting
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Optimizing the F-measure for Threshold-free Salient Object Detection
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
- Memory Optimization for Deep Networks
- SegSort: Segmentation by Discriminative Sorting of Segments
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Conditional Driving from Natural Language Instructions
- LeYOLO, New Embedded Architecture for Object Detection
- AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
- DHA: End-to-End Joint Optimization of Data Augmentation Policy, Hyper-parameter and Architecture
- Deep CNNs for Peripheral Blood Cell Classification
- Re-identification = Retrieval + Verification: Back to Essence and Forward with a New Metric
- Simultaneously Optimizing Weight and Quantizer of Ternary Neural Network using Truncated Gaussian Approximation
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- NaturalInversion: Data-Free Image Synthesis Improving Real-World Consistency
- Deep Learning for Efficient Reconstruction of High-Resolution Turbulent DNS Data
- Recurrent Neural Networks for video object detection
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Learning from scarce information: using synthetic data to classify Roman fine ware pottery
- Multi-objective Neural Architecture Search via Non-stationary Policy Gradient
- Vision and Tactile Robotic System to Grasp Litter in Outdoor Environments
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- Low-Power Computer Vision: Status, Challenges, Opportunities
- A Simple Method to Reduce Off-chip Memory Accesses on Convolutional Neural Networks
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- Intra-Ensemble in Neural Networks
- m-RevNet: Deep Reversible Neural Networks with Momentum
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Low-Cost Parameterizations of Deep Convolutional Neural Networks
- Approximations in Deep Learning
- DC-NAS: Divide-and-Conquer Neural Architecture Search
- Discovering Multi-Hardware Mobile Models via Architecture Search
- Dynamic-OFA: Runtime DNN Architecture Switching for Performance Scaling on Heterogeneous Embedded Platforms
- A Survey of Mobile Computing for the Visually Impaired
- Deep Transfer Learning for Automated Diagnosis of Skin Lesions from Photographs
- Matrix and tensor decompositions for training binary neural networks
- FedSC: Federated Learning with Semantic-Aware Collaboration
- Physics-aware Spatiotemporal Modules with Auxiliary Tasks for Meta-Learning
- ExplainFix: Explainable Spatially Fixed Deep Networks
- n-hot: Efficient bit-level sparsity for powers-of-two neural network quantization
- GAN-Knowledge Distillation for one-stage Object Detection
- EPNAS: Efficient Progressive Neural Architecture Search
- Real-Time Freespace Segmentation on Autonomous Robots for Detection of Obstacles and Drop-Offs
- How Does Supernet Help in Neural Architecture Search?
- WaveletNet: Logarithmic Scale Efficient Convolutional Neural Networks for Edge Devices
- Developing efficient transfer learning strategies for robust scene recognition in mobile robotics using pre-trained convolutional neural networks
- VTLayout: Fusion of Visual and Text Features for Document Layout Analysis
- Deep Learning for HDR Imaging: State-of-the-Art and Future Trends
- SMU: smooth activation function for deep networks using smoothing maximum technique
- MobiVSR: A Visual Speech Recognition Solution for Mobile Devices
- Diffusion models for Handwriting Generation
- Performance Evaluation of Convolutional Neural Networks for Gait Recognition
- Prioritized Architecture Sampling with Monto-Carlo Tree Search
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- HAO: Hardware-aware neural Architecture Optimization for Efficient Inference
- Evaluating Performance of an Adult Pornography Classifier for Child Sexual Abuse Detection
- Learning from a Lightweight Teacher for Efficient Knowledge Distillation
- What Deep CNNs Benefit from Global Covariance Pooling: An Optimization Perspective
- Unstructured Road Vanishing Point Detection Using the Convolutional Neural Network and Heatmap Regression
- A Light-Weighted Convolutional Neural Network for Bitemporal SAR Image Change Detection
- Supervised dimensionality reduction by a Linear Discriminant Analysis on pre-trained CNN features
- HBONet: Harmonious Bottleneck on Two Orthogonal Dimensions
- Instance Scale Normalization for image understanding
- LiDAR ICPS-net: Indoor Camera Positioning based-on Generative Adversarial Network for RGB to Point-Cloud Translation
- Geometry-constrained Car Recognition Using a 3D Perspective Network
- Dynamic Spatial Verification for Large-Scale Object-Level Image Retrieval
- A 3D CNN Network with BERT For Automatic COVID-19 Diagnosis From CT-Scan Images
- Dense xUnit Networks
- Greedy Network Enlarging
- APNN-TC: Accelerating Arbitrary Precision Neural Networks on Ampere GPU Tensor Cores
- Cross-Channel Intragroup Sparsity Neural Network
- Search Spaces for Neural Model Training
- You Only Search Once: A Fast Automation Framework for Single-Stage DNN/Accelerator Co-design
- ZeBRA: Precisely Destroying Neural Networks with Zero-Data Based Repeated Bit Flip Attack
- Manifestation of Image Contrast in Deep Networks
- AirPen: A Touchless Fingertip Based Gestural Interface for Smartphones and Head-Mounted Devices
- Integrating Multiple Receptive Fields through Grouped Active Convolution
- GCF-Net: Gated Clip Fusion Network for Video Action Recognition
- Full-Stack Filters to Build Minimum Viable CNNs
- HSCoNAS: Hardware-Software Co-Design of Efficient DNNs via Neural Architecture Search
- BN-NAS: Neural Architecture Search with Batch Normalization
- Understanding the Disharmony between Weight Normalization Family and Weight Decay: shifted Regularizer
- Learning Purified Feature Representations from Task-irrelevant Labels
- Multigrid-in-Channels Architectures for Wide Convolutional Neural Networks
- A New Training Framework for Deep Neural Network
- Multiple Myeloma Cancer Cell Instance Segmentation
- Learning to Cascade: Confidence Calibration for Improving the Accuracy and Computational Cost of Cascade Inference Systems
- Compact CNN Structure Learning by Knowledge Distillation
- MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
- Generating Adversarial Inputs Using A Black-box Differential Technique
- Learning and Exploiting Interclass Visual Correlations for Medical Image Classification
- TSDM: Tracking by SiamRPN++ with a Depth-refiner and a Mask-generator
- One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Reason induced visual attention for explainable autonomous driving
- Confidence Contours: Uncertainty-Aware Annotation for Medical Semantic Segmentation
- Mutually-aware Sub-Graphs Differentiable Architecture Search
- SVMax: A Feature Embedding Regularizer
- Can Targeted Adversarial Examples Transfer When the Source and Target Models Have No Label Space Overlap?
- Single Image Depth Prediction with Wavelet Decomposition
- Population Gradients improve performance across data-sets and architectures in object classification
- RingCNN: Exploiting Algebraically-Sparse Ring Tensors for Energy-Efficient CNN-Based Computational Imaging
- Selective Output Smoothing Regularization: Regularize Neural Networks by Softening Output Distributions
- DSXplore: Optimizing Convolutional Neural Networks via Sliding-Channel Convolutions
- One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space Shrinking
- SparseMask: Differentiable Connectivity Learning for Dense Image Prediction
- Safety Metrics for Semantic Segmentation in Autonomous Driving
- Distributed Layer-Partitioned Training for Privacy-Preserved Deep Learning
- Perceptron Synthesis Network: Rethinking the Action Scale Variances in Videos
- Query-Free Adversarial Transfer via Undertrained Surrogates
- Differentiable Neural Architecture Transformation for Reproducible Architecture Improvement
- Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural Networks
- InFL-UX: A Toolkit for Web-Based Interactive Federated Learning
- An Improved Relevance Feedback in CBIR
- An approach to hummed-tune and song sequences matching
- Multilayer Dense Connections for Hierarchical Concept Classification
- Learn to synthesize and synthesize to learn
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- Recognition Oriented Iris Image Quality Assessment in the Feature Space
- Fetal MRI by robust deep generative prior reconstruction and diffeomorphic registration: application to gestational age prediction
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- Fine-grained Optimization of Deep Neural Networks
- A CNN Accelerator on FPGA Using Depthwise Separable Convolution
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click Annotations
- Radius Adaptive Convolutional Neural Network
- Towards More Efficient and Effective Inference: The Joint Decision of Multi-Participants
- Self-Supervised Visual Representation Learning Using Lightweight Architectures
- DAC: Data-free Automatic Acceleration of Convolutional Networks
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks
- Against Membership Inference Attack: Pruning is All You Need