Going Deeper with Convolutions
arXiv:1409.4842
Abstract
We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC 2014). The main hallmark of this architecture is the improved utilization of the computing resources inside the network. This was achieved by a carefully crafted design that allows for increasing the depth and width of the network while keeping the computational budget constant. To optimize quality, the architectural decisions were based on the Hebbian principle and the intuition of multi-scale processing. One particular incarnation used in our submission for ILSVRC 2014 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection.
Cited by in corpus (44)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Explaining and Harnessing Adversarial Examples
- Striving for Simplicity: The All Convolutional Net
- FitNets: Hints for Thin Deep Nets
- Learning Face Representation from Scratch
- Compressing Deep Convolutional Networks using Vector Quantization
- DeepID3: Face Recognition with Very Deep Neural Networks
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Learning Deconvolution Network for Semantic Segmentation
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- An Empirical Evaluation of Deep Learning on Highway Driving
- Massively Multitask Networks for Drug Discovery
- Fractional Max-Pooling
- Beyond Short Snippets: Deep Networks for Video Classification
- Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
- Training Deeper Convolutional Networks with Deep Supervision
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Memory Bounded Deep Convolutional Networks
- Attention for Fine-Grained Categorization
- Evaluating Two-Stream CNN for Video Classification
- APAC: Augmented PAttern Classification with Neural Networks
- Scale-Invariant Convolutional Neural Networks
- Exploiting Local Features from Deep Networks for Image Retrieval
- Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- A Probabilistic Theory of Deep Learning
- segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection
- Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image Interpretation
- Deep Spatial Pyramid: The Devil is Once Again in the Details
- Training Binary Multilayer Neural Networks for Image Classification using Expectation Backpropagation
- When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Convolutional Neural Networks at Constrained Time Cost
- Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Material Recognition in the Wild with the Materials in Context Database
- Boosting Convolutional Features for Robust Object Proposals
- Deep Learning for Object Saliency Detection and Image Segmentation
- Object-centric Sampling for Fine-grained Image Classification
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Generative Class-conditional Autoencoders
- Purine: A bi-graph based deep learning framework
- Improving Image Classification with Location Context