OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
arXiv:1312.6229
Abstract
We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be efficiently implemented within a ConvNet. We also introduce a novel deep learning approach to localization by learning to predict object boundaries. Bounding boxes are then accumulated rather than suppressed in order to increase detection confidence. We show that different tasks can be learned simultaneously using a single shared network. This integrated framework is the winner of the localization task of the ImageNet Large Scale Visual Recognition Challenge 2013 (ILSVRC2013) and obtained very competitive results for the detection and classifications tasks. In post-competition work, we establish a new state of the art for the detection task. Finally, we release a feature extractor from our best model called OverFeat.
References in corpus (1)
Cited by in corpus (177)
- Deep Learning in Neural Networks: An Overview
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Deep Domain Confusion: Maximizing for Domain Invariance
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
- Going Deeper with Convolutions
- Recent advances and clinical applications of deep learning in medical image analysis
- cuDNN: Efficient Primitives for Deep Learning
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Road Damage Detection Using Deep Neural Networks with Images Captured Through a Smartphone
- Segmentation-Based Deep-Learning Approach for Surface-Defect Detection
- MeshCNN: A Network with an Edge
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Fixed Point Quantization of Deep Convolutional Networks
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Visualizing Deep Convolutional Neural Networks Using Natural Pre-Images
- P-wave arrival picking and first-motion polarity determination with deep learning
- Medical Image Retrieval using Deep Convolutional Neural Network
- PlaNet - Photo Geolocation with Convolutional Neural Networks
- Residual Networks of Residual Networks: Multilevel Residual Networks
- An In-field Automatic Wheat Disease Diagnosis System
- Recent Advances in Convolutional Neural Networks
- Towards automatic pulmonary nodule management in lung cancer screening with deep learning
- Domain Adaptation for Visual Applications: A Comprehensive Survey
- Fast Automated Analysis of Strong Gravitational Lenses with Convolutional Neural Networks
- Convolutional Neural Network-based Place Recognition
- Regularization for Deep Learning: A Taxonomy
- Visual Identification of Individual Holstein-Friesian Cattle via Deep Metric Learning
- RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
- End-to-End Photo-Sketch Generation via Fully Convolutional Representation Learning
- VPR-Bench: An Open-Source Visual Place Recognition Evaluation Framework with Quantifiable Viewpoint and Appearance Change
- ImageNet pre-trained models with batch normalization
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- Memory Bounded Deep Convolutional Networks
- Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
- The Origins and Prevalence of Texture Bias in Convolutional Neural Networks
- Deep Residual Networks with Exponential Linear Unit
- Scale-aware Fast R-CNN for Pedestrian Detection
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- Physical Adversarial Examples for Object Detectors
- Weakly Supervised Medical Diagnosis and Localization from Multiple Resolutions
- Exploiting Local Features from Deep Networks for Image Retrieval
- A Computer Vision System to Localize and Classify Wastes on the Streets
- Feature Representation in Convolutional Neural Networks
- DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection
- Object Detection Through Exploration With A Foveated Visual Field
- NeuralPower: Predict and Deploy Energy-Efficient Convolutional Neural Networks
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Machine-learning techniques for fast and accurate feature localization in holograms of colloidal particles
- Deep Convolution Networks for Compression Artifacts Reduction
- VddNet: Vine Disease Detection Network Based on Multispectral Images and Depth Map
- SDCT-AuxNet: DCT Augmented Stain Deconvolutional CNN with Auxiliary Classifier for Cancer Diagnosis
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Learning Modulated Loss for Rotated Object Detection
- Deep learning with Elastic Averaging SGD
- MaskConnect: Connectivity Learning by Gradient Descent
- Deformable Part Models are Convolutional Neural Networks
- Local Decorrelation For Improved Detection
- Synthesized Classifiers for Zero-Shot Learning
- Overcoming Small Minirhizotron Datasets Using Transfer Learning
- You Only Hear Once: A YOLO-like Algorithm for Audio Segmentation and Sound Event Detection
- Fine-grained pose prediction, normalization, and recognition
- MEC: Memory-efficient Convolution for Deep Neural Network
- The use of convolutional neural networks for modelling large optically-selected strong galaxy-lens samples
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- Data Distillation: Towards Omni-Supervised Learning
- Visual Affordance and Function Understanding: A Survey
- The Lovász-Softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks
- Learning non-maximum suppression
- In-domain representation learning for remote sensing
- Monocular Object Instance Segmentation and Depth Ordering with CNNs
- Workpiece Image-based Tool Wear Classification in Blanking Processes Using Deep Convolutional Neural Networks
- Deep Convolutional Features for Image Based Retrieval and Scene Categorization
- End-to-end people detection in crowded scenes
- Learning Sensor Multiplexing Design through Back-propagation
- MegDet: A Large Mini-Batch Object Detector
- Automated Architecture Design for Deep Neural Networks
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Maximum-Margin Structured Learning with Deep Networks for 3D Human Pose Estimation
- Learning Deep Object Detectors from 3D Models
- Understanding Deep Architectures using a Recursive Convolutional Network
- Connectivity Learning in Multi-Branch Networks
- Perfect density models cannot guarantee anomaly detection
- Compact Convolutional Neural Network Cascade for Face Detection
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
- Side-Aware Boundary Localization for More Precise Object Detection
- Deep Learning for Object Saliency Detection and Image Segmentation
- Virtual Training for a Real Application: Accurate Object-Robot Relative Localization without Calibration
- MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-based Protein Structure Prediction
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
- A Fast Face Detection Method via Convolutional Neural Network
- Fingertip in the Eye: A cascaded CNN pipeline for the real-time fingertip detection in egocentric videos
- Digging Deep into the layers of CNNs: In Search of How CNNs Achieve View Invariance
- Deep Representation of Facial Geometric and Photometric Attributes for Automatic 3D Facial Expression Recognition
- Fast Fourier Transformation for Optimizing Convolutional Neural Networks in Object Recognition
- From Image-level to Pixel-level Labeling with Convolutional Networks
- Generic Object Detection With Dense Neural Patterns and Regionlets
- cltorch: a Hardware-Agnostic Backend for the Torch Deep Neural Network Library, Based on OpenCL
- Convolutional Models for Joint Object Categorization and Pose Estimation
- Learning Structured Inference Neural Networks with Label Relations
- WordFence: Text Detection in Natural Images with Border Awareness
- Exploring the Applications of Faster R-CNN and Single-Shot Multi-box Detection in a Smart Nursery Domain
- Advances in Deep Learning for Hyperspectral Image Analysis--Addressing Challenges Arising in Practical Imaging Scenarios
- Self-produced Guidance for Weakly-supervised Object Localization
- SA-CNN: Dynamic Scene Classification using Convolutional Neural Networks
- ScratchDet: Training Single-Shot Object Detectors from Scratch
- Comparing Apples and Oranges: Off-Road Pedestrian Detection on the NREC Agricultural Person-Detection Dataset
- Identifying Land Patterns from Satellite Imagery in Amazon Rainforest using Deep Learning
- Window-Object Relationship Guided Representation Learning for Generic Object Detections
- Learning Fine-grained Features via a CNN Tree for Large-scale Classification
- Detector Discovery in the Wild: Joint Multiple Instance and Representation Learning
- Deep Regionlets for Object Detection
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- Supervised and Unsupervised End-to-End Deep Learning for Gene Ontology Classification of Neural In Situ Hybridization Images
- Exploit Bounding Box Annotations for Multi-label Object Recognition
- Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
- Novelty Detection in MultiClass Scenarios with Incomplete Set of Class Labels
- Semantic Bilinear Pooling for Fine-Grained Recognition
- Seeing What Is Not There: Learning Context to Determine Where Objects Are Missing
- Unsupervised Regenerative Learning of Hierarchical Features in Spiking Deep Networks for Object Recognition
- Brain Abnormality Detection by Deep Convolutional Neural Network
- Part-based R-CNNs for Fine-grained Category Detection
- Adaptive Weight Assignment Scheme For Multi-task Learning
- Hierarchical Bayesian Noise Inference for Robust Real-time Probabilistic Object Classification
- Parsing Occluded People by Flexible Compositions
- Jet Single Shot Detection
- Material Classification in the Wild: Do Synthesized Training Data Generalise Better than Real-World Training Data?
- Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
- Learning from Web Data: the Benefit of Unsupervised Object Localization
- Fast Object Localization Using a CNN Feature Map Based Multi-Scale Search
- Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
- CNN based texture synthesize with Semantic segment
- Integrated Inference and Learning of Neural Factors in Structural Support Vector Machines
- End-to-End Integration of a Convolutional Network, Deformable Parts Model and Non-Maximum Suppression
- Universum Prescription: Regularization using Unlabeled Data
- Learning Image Conditioned Label Space for Multilabel Classification
- Design of Kernels in Convolutional Neural Networks for Image Classification
- An Overview on Data Representation Learning: From Traditional Feature Learning to Recent Deep Learning
- Master's Thesis : Deep Learning for Visual Recognition
- Learning to Select Pre-Trained Deep Representations with Bayesian Evidence Framework
- Character-Based Text Classification using Top Down Semantic Model for Sentence Representation
- Localized Traffic Sign Detection with Multi-scale Deconvolution Networks
- Multi-label Pixelwise Classification for Reconstruction of Large-scale Urban Areas
- Optimising the Input Image to Improve Visual Relationship Detection
- Saliency Driven Object recognition in egocentric videos with deep CNN
- An efficient deep learning hashing neural network for mobile visual search
- Learning to Navigate for Fine-grained Classification
- Progressive Representation Adaptation for Weakly Supervised Object Localization
- Convolution in Convolution for Network in Network
- Natural Adversarial Objects
- Image Classification with A Deep Network Model based on Compressive Sensing
- Better Exploiting OS-CNNs for Better Event Recognition in Images
- Robust Optimization for Deep Regression
- Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object Detection
- The role of a layer in deep neural networks: a Gaussian Process perspective
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- Molecular Sparse Representation by 3D Ellipsoid Radial Basis Function Neural Networks via Regularization
- Investigations of the Influences of a CNN's Receptive Field on Segmentation of Subnuclei of Bilateral Amygdalae
- Compact retail shelf segmentation for mobile deployment
- Learning to Generate Content-Aware Dynamic Detectors
- DropRegion Training of Inception Font Network for High-Performance Chinese Font Recognition
- Deep Discriminative Model for Video Classification
- Efficient Modelling Across Time of Human Actions and Interactions
- PolyScientist: Automatic Loop Transformations Combined with Microkernels for Optimization of Deep Learning Primitives
- Active Object Localization with Deep Reinforcement Learning
- Compact Deep Aggregation for Set Retrieval
- DeepKey: Towards End-to-End Physical Key Replication From a Single Photograph
- Automatic Inspection of Utility Scale Solar Power Plants using Deep Learning
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- FHEDN: A based on context modeling Feature Hierarchy Encoder-Decoder Network for face detection
- Compression Artifacts Reduction by a Deep Convolutional Network
- Ventral-Dorsal Neural Networks: Object Detection via Selective Attention
- A Hybrid Framework for Matching Printing Design Files to Product Photos
- Primary Tumor Origin Classification of Lung Nodules in Spectral CT using Transfer Learning