Rethinking the Inception Architecture for Computer Vision
arXiv:1512.00567
Abstract
Convolutional networks are at the core of most state-of-the-art computer vision solutions for a wide variety of tasks. Since 2014 very deep convolutional networks started to become mainstream, yielding substantial gains in various benchmarks. Although increased model size and computational cost tend to translate to immediate quality gains for most tasks (as long as enough labeled data is provided for training), computational efficiency and low parameter count are still enabling factors for various use cases such as mobile vision and big-data scenarios. Here we explore ways to scale up networks in ways that aim at utilizing the added computation as efficiently as possible by suitably factorized convolutions and aggressive regularization. We benchmark our methods on the ILSVRC 2012 classification challenge validation set demonstrate substantial gains over the state of the art: 21.2% top-1 and 5.6% top-5 error for single frame evaluation using a network with a computational cost of 5 billion multiply-adds per inference and with using less than 25 million parameters. With an ensemble of 4 models and multi-crop evaluation, we report 3.5% top-5 error on the validation set (3.6% error on the test set) and 17.3% top-1 error on the validation set.
Cited by in corpus (168)
- TensorFlow: A system for large-scale machine learning
- Diffusion Models Beat GANs on Image Synthesis
- Conditional Image Synthesis With Auxiliary Classifier GANs
- Densely Connected Convolutional Networks
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Robust Physical-World Attacks on Deep Learning Models
- 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation
- Revisiting Batch Normalization For Practical Domain Adaptation
- Black-box Adversarial Attacks with Limited Queries and Information
- Satellite Pose Estimation Challenge: Dataset, Competition Design and Results
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- Deep Learning in the Automotive Industry: Applications and Tools
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Hydra: an Ensemble of Convolutional Neural Networks for Geospatial Land Classification
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Faster CryptoNets: Leveraging Sparsity for Real-World Encrypted Inference
- A Pursuit of Temporal Accuracy in General Activity Detection
- Imagination improves Multimodal Translation
- Variants of RMSProp and Adagrad with Logarithmic Regret Bounds
- Transfer Learning using CNN for Handwritten Devanagari Character Recognition
- Re-ID done right: towards good practices for person re-identification
- Domain Adaptive Transfer Learning with Specialist Models
- Deep Koalarization: Image Colorization using CNNs and Inception-ResNet-v2
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Interleaved Group Convolutions for Deep Neural Networks
- The Devil is in the Middle: Exploiting Mid-level Representations for Cross-Domain Instance Matching
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Vision Transformer based COVID-19 Detection using Chest X-rays
- Co-training for Demographic Classification Using Deep Learning from Label Proportions
- Learning Feature Pyramids for Human Pose Estimation
- Two-Stream Neural Networks for Tampered Face Detection
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- Multiscale Vision Transformers
- Boosting Adversarial Attacks with Momentum
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Prospects for Theranostics in Neurosurgical Imaging: Empowering Confocal Laser Endomicroscopy Diagnostics via Deep Learning
- Gmail Smart Compose: Real-Time Assisted Writing
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- An Empirical Study on Writer Identification & Verification from Intra-variable Individual Handwriting
- Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis
- The LHC Olympics 2020: A Community Challenge for Anomaly Detection in High Energy Physics
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Local minima in training of neural networks
- A global method to identify trees outside of closed-canopy forests with medium-resolution satellite imagery
- Multi-scale Deep Learning Architectures for Person Re-identification
- Camera Style Adaptation for Person Re-identification
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- SFD: Single Shot Scale-invariant Face Detector
- Machine learning for complete intersection Calabi-Yau manifolds: a methodological study
- PDNet: Semantic Segmentation integrated with a Primal-Dual Network for Document binarization
- Deep Pyramidal Residual Networks
- Improving Object Counting with Heatmap Regulation
- Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Neural Networks with Recurrent Generative Feedback
- Query-Efficient Black-box Adversarial Examples (superceded)
- Histopathologic Image Processing: A Review
- Quantization for Rapid Deployment of Deep Neural Networks
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- Global Encoding for Abstractive Summarization
- Practical Block-wise Neural Network Architecture Generation
- A survey of Object Classification and Detection based on 2D/3D data
- A Multi-Stream Convolutional Neural Network Framework for Group Activity Recognition
- Graph-Structured Visual Imitation
- PredNet and Predictive Coding: A Critical Review
- From Third Person to First Person: Dataset and Baselines for Synthesis and Retrieval
- Learning to Train a Binary Neural Network
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- The iNaturalist Species Classification and Detection Dataset
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- Cell nuclei classification in histopathological images using hybrid OLConvNet
- DOOM Level Generation using Generative Adversarial Networks
- A deep learning based solution for construction equipment detection: from development to deployment
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- End-to-end Video-level Representation Learning for Action Recognition
- CNN-based Facial Affect Analysis on Mobile Devices
- Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection
- Differentiable Architecture Search with Ensemble Gumbel-Softmax
- Deep Neural Network Concepts for Background Subtraction: A Systematic Review and Comparative Evaluation
- Dual Path Networks for Multi-Person Human Pose Estimation
- PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
- Realistic Image Generation using Region-phrase Attention
- IMG2SMI: Translating Molecular Structure Images to Simplified Molecular-input Line-entry System
- Compact Global Descriptor for Neural Networks
- Audio-video Emotion Recognition in the Wild using Deep Hybrid Networks
- FireNet: Real-time Segmentation of Fire Perimeter from Aerial Video
- Benanza: Automatic Benchmark Generation to Compute "Lower-bound" Latency and Inform Optimizations of Deep Learning Models on GPUs
- Multimodal Memory Modelling for Video Captioning
- LaSO: Label-Set Operations networks for multi-label few-shot learning
- Frequency Centric Defense Mechanisms against Adversarial Examples
- Me, Myself and My Killfie: Characterizing and Preventing Selfie Deaths
- Statistically Motivated Second Order Pooling
- Evaluation of Deep Species Distribution Models using Environment and Co-occurrences
- BrainSlug: Transparent Acceleration of Deep Learning Through Depth-First Parallelism
- Designing a Micro-Benchmark Suite to Evaluate gRPC for TensorFlow: Early Experiences
- Top-down Visual Saliency Guided by Captions
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- Comparing Deep Learning Models for Multi-cell Classification in Liquid-based Cervical Cytology Images
- Optimizing Prediction Serving on Low-Latency Serverless Dataflow
- Feature based Sequential Classifier with Attention Mechanism
- Predicting molecular phenotypes from histopathology images: a transcriptome-wide expression-morphology analysis in breast cancer
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Learning 3D Shapes as Multi-Layered Height-maps using 2D Convolutional Networks
- A Large-scale Attribute Dataset for Zero-shot Learning
- Tag Prediction at Flickr: a View from the Darkroom
- Chest X-Rays Image Classification from beta-Variational Autoencoders Latent Features
- ComicGAN: Text-to-Comic Generative Adversarial Network
- Irregular Convolutional Neural Networks
- Large-Scale 3D Scene Classification With Multi-View Volumetric CNN
- DeepFolio: Convolutional Neural Networks for Portfolios with Limit Order Book Data
- Design of Efficient Deep Learning models for Determining Road Surface Condition from Roadside Camera Images and Weather Data
- Dual Encoder Fusion U-Net (DEFU-Net) for Cross-manufacturer Chest X-ray Segmentation
- Taming GANs with Lookahead-Minmax
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- MMGAN: Manifold Matching Generative Adversarial Network
- Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
- Less is More: Sparse Sampling for Dense Reaction Predictions
- Modular Learning Component Attacks: Today's Reality, Tomorrow's Challenge
- Reconstruction of Simulation-Based Physical Field by Reconstruction Neural Network Method
- Winter Road Surface Condition Recognition Using A Pretrained Deep Convolutional Network
- Rep Works in Speaker Verification
- WaveletNet: Logarithmic Scale Efficient Convolutional Neural Networks for Edge Devices
- Bio-Measurements Estimation and Support in Knee Recovery through Machine Learning
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- RPC Considered Harmful: Fast Distributed Deep Learning on RDMA
- Chained Predictions Using Convolutional Neural Networks
- ModiPick: SLA-aware Accuracy Optimization For Mobile Deep Inference
- Learning scale-variant features for robust iris authentication with deep learning based ensemble framework
- SuperNet -- An efficient method of neural networks ensembling
- Ensemble Transfer Learning for Emergency Landing Field Identification on Moderate Resource Heterogeneous Kubernetes Cluster
- Visual Themes and Sentiment on Social Networks To Aid First Responders During Crisis Events
- Compression Fractures Detection on CT
- Zoom-RNN: A Novel Method for Person Recognition Using Recurrent Neural Networks
- Semantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
- Image-based Vehicle Re-identification Model with Adaptive Attention Modules and Metadata Re-ranking
- Manifestation of Image Contrast in Deep Networks
- A comparison of deep machine learning algorithms in COVID-19 disease diagnosis
- Human Activity Recognition for Edge Devices
- Downscaling Attack and Defense: Turning What You See Back Into What You Get
- Multi-vision Attention Networks for On-line Red Jujube Grading
- Augmentation Inside the Network
- Sketches image analysis: Web image search engine usingLSH index and DNN InceptionV3
- Classifying logistic vehicles in cities using Deep learning
- Cascading Neural Network Methodology for Artificial Intelligence-Assisted Radiographic Detection and Classification of Lead-Less Implanted Electronic Devices within the Chest
- A Multi-Scale CNN and Curriculum Learning Strategy for Mammogram Classification
- An efficient deep learning hashing neural network for mobile visual search
- Toward Accurate Platform-Aware Performance Modeling for Deep Neural Networks
- Large-scale mammography CAD with Deformable Conv-Nets
- Synthetic Generation of Three-Dimensional Cancer Cell Models from Histopathological Images
- Caramel: Accelerating Decentralized Distributed Deep Learning with Computation Scheduling
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- DeepWheat: Estimating Phenotypic Traits from Crop Images with Deep Learning
- A Computer Vision Approach to Combat Lyme Disease
- Localized Adversarial Training for Increased Accuracy and Robustness in Image Classification
- Attack Transferability Characterization for Adversarially Robust Multi-label Classification
- User-centric Composable Services: A New Generation of Personal Data Analytics
- FSD: Feature Skyscraper Detector for Stem End and Blossom End of Navel Orange
- Leveraging Model Interpretability and Stability to increase Model Robustness
- Regularized adversarial examples for model interpretability
- The Focus-Aspect-Polarity Model for Predicting Subjective Noun Attributes in Images
- Pay Attention to Convolution Filters: Towards Fast and Accurate Fine-Grained Transfer Learning