Xception: Deep Learning with Depthwise Separable Convolutions
arXiv:1610.02357
Abstract
We present an interpretation of Inception modules in convolutional neural networks as being an intermediate step in-between regular convolution and the depthwise separable convolution operation (a depthwise convolution followed by a pointwise convolution). In this light, a depthwise separable convolution can be understood as an Inception module with a maximally large number of towers. This observation leads us to propose a novel deep convolutional neural network architecture inspired by Inception, where Inception modules have been replaced with depthwise separable convolutions. We show that this architecture, dubbed Xception, slightly outperforms Inception V3 on the ImageNet dataset (which Inception V3 was designed for), and significantly outperforms Inception V3 on a larger image classification dataset comprising 350 million images and 17,000 classes. Since the Xception architecture has the same number of parameters as Inception V3, the performance gains are not due to increased capacity but rather to a more efficient use of model parameters.
References in corpus (3)
Cited by in corpus (175)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Neural Architecture Search: A Survey
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- SchNet: A continuous-filter convolutional neural network for modeling quantum interactions
- AlignedReID: Surpassing Human-Level Performance in Person Re-Identification
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
- Fast-SCNN: Fast Semantic Segmentation Network
- Channel Pruning for Accelerating Very Deep Neural Networks
- Hello Edge: Keyword Spotting on Microcontrollers
- CBAM: Convolutional Block Attention Module
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Compact Convolutional Neural Networks for Classification of Asynchronous Steady-state Visual Evoked Potentials
- One Model To Learn Them All
- Depthwise Separable Convolutions for Neural Machine Translation
- Slimmable Neural Networks
- Hierarchical Representations for Efficient Architecture Search
- W-Net: A Deep Model for Fully Unsupervised Image Segmentation
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- ContextNet: Exploring Context and Detail for Semantic Segmentation in Real-time
- The Reversible Residual Network: Backpropagation Without Storing Activations
- L1-Norm Batch Normalization for Efficient Training of Deep Neural Networks
- CvT: Introducing Convolutions to Vision Transformers
- Hydra: an Ensemble of Convolutional Neural Networks for Geospatial Land Classification
- StegNet: Mega Image Steganography Capacity with Deep Convolutional Network
- Machine Learning for the Zwicky Transient Facility
- Using transfer learning to detect galaxy mergers
- XNOR Neural Engine: a Hardware Accelerator IP for 21.6 fJ/op Binary Neural Network Inference
- Margin Sample Mining Loss: A Deep Learning Based Method for Person Re-identification
- PointCNN: Convolution On -Transformed Points
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- 2018 Robotic Scene Segmentation Challenge
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- Factorization tricks for LSTM networks
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Interleaved Group Convolutions for Deep Neural Networks
- Network Decoupling: From Regular to Depthwise Separable Convolutions
- UnsuperPoint: End-to-end Unsupervised Interest Point Detector and Descriptor
- Towards Practical Verification of Machine Learning: The Case of Computer Vision Systems
- Co-training for Demographic Classification Using Deep Learning from Label Proportions
- Real-time Convolutional Neural Networks for Emotion and Gender Classification
- Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
- Driving Scene Perception Network: Real-time Joint Detection, Depth Estimation and Semantic Segmentation
- MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices
- Deep Convolutional Neural Network Design Patterns
- ShuffleSeg: Real-time Semantic Segmentation Network
- An Empirical Study on Writer Identification & Verification from Intra-variable Individual Handwriting
- Detecting solar system objects with convolutional neural networks
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Shift: A Zero FLOP, Zero Parameter Alternative to Spatial Convolutions
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Towards High Performance Video Object Detection for Mobiles
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- Instance Segmentation by Deep Coloring
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- IGCV: Interleaved Structured Sparse Convolutional Neural Networks
- Explaining deep learning of galaxy morphology with saliency mapping
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Quantization for Rapid Deployment of Deep Neural Networks
- Question-Guided Hybrid Convolution for Visual Question Answering
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Practical Block-wise Neural Network Architecture Generation
- Intelligence Beyond the Edge: Inference on Intermittent Embedded Systems
- Learning Chained Deep Features and Classifiers for Cascade in Object Detection
- Convolution, attention and structure embedding
- Learning Discriminative Motion Features Through Detection
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- BlockQNN: Efficient Block-wise Neural Network Architecture Generation
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Integral Human Pose Regression
- Measuring the Transferability of Adversarial Examples
- MirBot: A collaborative object recognition system for smartphones using convolutional neural networks
- FD-MobileNet: Improved MobileNet with a Fast Downsampling Strategy
- Tensor2Tensor for Neural Machine Translation
- An Experimental Study of the Impact of Pre-training on the Pruning of a Convolutional Neural Network
- Improving Image Clustering With Multiple Pretrained CNN Feature Extractors
- EffNet: An Efficient Structure for Convolutional Neural Networks
- Road Segmentation Using CNN with GRU
- Improved Regularization Techniques for End-to-End Speech Recognition
- Adversarial Alignment of Class Prediction Uncertainties for Domain Adaptation
- SAWNet: A Spatially Aware Deep Neural Network for 3D Point Cloud Processing
- Efficient Deep Neural Networks
- Convolutional Networks with Dense Connectivity
- SingleGAN: Image-to-Image Translation by a Single-Generator Network using Multiple Generative Adversarial Learning
- Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration
- Performance assessment of the deep learning technologies in grading glaucoma severity
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- Fast Neural Architecture Construction using EnvelopeNets
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Adaptive Weighting Multi-Field-of-View CNN for Semantic Segmentation in Pathology
- Gradient-based Training of Slow Feature Analysis by Differentiable Approximate Whitening
- A Note on Deepfake Detection with Low-Resources
- Improving End-to-End Speech Recognition with Policy Learning
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Making Neural Machine Reading Comprehension Faster
- Predicting bulge to total luminosity ratio of galaxies using deep learning
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- ESAI: Efficient Split Artificial Intelligence via Early Exiting Using Neural Architecture Search
- Efficient Smoothing of Dilated Convolutions for Image Segmentation
- Seesaw-Net: Convolution Neural Network With Uneven Group Convolution
- Towards High Performance Video Object Detection
- NeuNetS: An Automated Synthesis Engine for Neural Network Design
- Fire SSD: Wide Fire Modules based Single Shot Detector on Edge Device
- Estimating the Brittleness of AI: Safety Integrity Levels and the Need for Testing Out-Of-Distribution Performance
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Deep Image Orientation Angle Detection
- Deep Learning-based Aerial Image Segmentation with Open Data for Disaster Impact Assessment
- FMCode: A 3D In-the-Air Finger Motion Based User Login Framework for Gesture Interface
- Rethink ReLU to Training Better CNNs
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- On Machine Learning and Structure for Mobile Robots
- Multi-domain semantic segmentation with pyramidal fusion
- Design of Efficient Deep Learning models for Determining Road Surface Condition from Roadside Camera Images and Weather Data
- Fingerprint Spoof Buster
- Learning to Segment Human Body Parts with Synthetically Trained Deep Convolutional Networks
- Future Segmentation Using 3D Structure
- Classifying CMB time-ordered data through deep neural networks
- Ebola Optimization Search Algorithm (EOSA): A new metaheuristic algorithm based on the propagation model of Ebola virus disease
- ResNetX: a more disordered and deeper network architecture
- An Evolution of CNN Object Classifiers on Low-Resolution Images
- Merging and Evolution: Improving Convolutional Neural Networks for Mobile Applications
- WSNet: Compact and Efficient Networks Through Weight Sampling
- Towards Efficient Convolutional Neural Network for Domain-Specific Applications on FPGA
- Predicting Length of Stay in the Intensive Care Unit with Temporal Pointwise Convolutional Networks
- MOSQUITO-NET: A deep learning based CADx system for malaria diagnosis along with model interpretation using GradCam and class activation maps
- EPNAS: Efficient Progressive Neural Architecture Search
- High Performance Depthwise and Pointwise Convolutions on Mobile Devices
- WaveletNet: Logarithmic Scale Efficient Convolutional Neural Networks for Edge Devices
- SwiftSRGAN -- Rethinking Super-Resolution for Efficient and Real-time Inference
- One-class Steel Detector Using Patch GAN Discriminator for Visualising Anomalous Feature Map
- Model Optimization for Deep Space Exploration via Simulators and Deep Learning
- OMG - Emotion Challenge Solution
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Human Face Expressions from Images - 2D Face Geometry and 3D Face Local Motion versus Deep Neural Features
- GM-Net: Learning Features with More Efficiency
- CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
- Incomplete Dot Products for Dynamic Computation Scaling in Neural Network Inference
- Dense xUnit Networks
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- DeepMerge: Classifying High-redshift Merging Galaxies with Deep Neural Networks
- ESFNet: Efficient Network for Building Extraction from High-Resolution Aerial Images
- Sequential Random Network for Fine-grained Image Classification
- CondenseNet V2: Sparse Feature Reactivation for Deep Networks
- Road Segmentation Using CNN and Distributed LSTM
- The Ethical Dilemma when (not) Setting up Cost-based Decision Rules in Semantic Segmentation
- Identification of images of COVID-19 from Chest Computed Tomography (CT) images using Deep learning: Comparing COGNEX VisionPro Deep Learning 1.0 Software with Open Source Convolutional Neural Networks
- Lenses In VoicE (LIVE): Searching for strong gravitational lenses in the VOICE@VST survey using Convolutional Neural Networks
- FPAN: Fine-grained and Progressive Attention Localization Network for Data Retrieval
- Multi-Class Zero-Shot Learning for Artistic Material Recognition
- Augmentation Inside the Network
- Prune2Edge: A Multi-Phase Pruning Pipelines to Deep Ensemble Learning in IIoT
- Quantization Mimic: Towards Very Tiny CNN for Object Detection
- cofga: A Dataset for Fine Grained Classification of Objects from Aerial Imagery
- Facial Expressions Recognition with Convolutional Neural Networks
- Optimal Transfer Learning Model for Binary Classification of Funduscopic Images through Simple Heuristics
- Crowd-Sourced Road Quality Mapping in the Developing World
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Machine Learning and Deep Learning methods for predictive modelling from Raman spectra in bioprocessing
- Road Mapping in Low Data Environments with OpenStreetMap
- Domain Adaptation with Morphologic Segmentation
- DCNNs: A Transfer Learning comparison of Full Weapon Family threat detection for Dual-Energy X-Ray Baggage Imagery
- Bit-level Optimized Neural Network for Multi-antenna Channel Quantization
- Towards Solving the DeepFake Problem : An Analysis on Improving DeepFake Detection using Dynamic Face Augmentation
- PolyScientist: Automatic Loop Transformations Combined with Microkernels for Optimization of Deep Learning Primitives
- On evaluating CNN representations for low resource medical image classification
- Full-stack Optimization for Accelerating CNNs with FPGA Validation
- Deep Single Image Deraining Via Estimating Transmission and Atmospheric Light in rainy Scenes
- Remote Estimation of Free-Flow Speeds
- A CNN Accelerator on FPGA Using Depthwise Separable Convolution
- Classification of COVID-19 from CXR Images in a 15-class Scenario: an Attempt to Avoid Bias in the System
- Real-Time Facial Expression Emoji Masking with Convolutional Neural Networks and Homography
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification