CNN Features off-the-shelf: an Astounding Baseline for Recognition
arXiv:1403.6382
Abstract
Recent results indicate that the generic descriptors extracted from the convolutional neural networks are very powerful. This paper adds to the mounting evidence that this is indeed the case. We report on a series of experiments conducted for different recognition tasks using the publicly available code and model of the \overfeat network which was trained to perform object classification on ILSVRC13. We use features extracted from the \overfeat network as a generic image representation to tackle the diverse range of recognition tasks of object image classification, scene recognition, fine grained recognition, attribute detection and image retrieval applied to a diverse set of datasets. We selected these tasks and datasets as they gradually move further away from the original task and data the \overfeat network was trained to solve. Astonishingly, we report consistent superior results compared to the highly tuned state-of-the-art systems in all the visual classification tasks on various datasets. For instance retrieval it consistently outperforms low memory footprint methods except for sculptures dataset. The results are achieved using a linear SVM classifier (or distance in case of retrieval) applied to a feature representation of size 4096 extracted from a layer in the net. The representations are further modified using simple augmentation techniques e.g. jittering. The results strongly suggest that features obtained from deep learning with convolutional nets should be the primary candidate in most visual recognition tasks.
version 3 revisions: 1)Added results using feature processing and data augmentation 2)Referring to most recent efforts of using CNN for different visual recognition tasks 3) updated text/caption
Cited by in corpus (241)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?
- Deep Facial Expression Recognition: A Survey
- Recent Trends in Deep Learning Based Natural Language Processing
- Compressing Deep Convolutional Networks using Vector Quantization
- Meta-SGD: Learning to Learn Quickly for Few-Shot Learning
- Adversarial Feature Learning
- Particular object retrieval with integral max-pooling of CNN activations
- Fixed Point Quantization of Deep Convolutional Networks
- Land Use Classification in Remote Sensing Images by Convolutional Neural Networks
- PlaNet - Photo Geolocation with Convolutional Neural Networks
- Object-Part Attention Model for Fine-grained Image Classification
- Learning Features for Offline Handwritten Signature Verification using Deep Convolutional Neural Networks
- Deep Image: Scaling up Image Recognition
- What makes ImageNet good for transfer learning?
- Multi-view Convolutional Neural Networks for 3D Shape Recognition
- Learning Deep Embeddings with Histogram Loss
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- A Comprehensive Study of Class Incremental Learning Algorithms for Visual Tasks
- CNN-RNN: A Unified Framework for Multi-label Image Classification
- Convolutional Neural Network-based Place Recognition
- Learning to Compare Image Patches via Convolutional Neural Networks
- A Richly Annotated Dataset for Pedestrian Attribute Recognition
- Pose Invariant Embedding for Deep Person Re-identification
- Recent Advance in Content-based Image Retrieval: A Literature Survey
- Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- The Devil is in the Tails: Fine-grained Classification in the Wild
- Good Practice in CNN Feature Transfer
- Fully Convolutional Neural Networks for Crowd Segmentation
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring
- Universal representations:The missing link between faces, text, planktons, and cat breeds
- Imbalanced Malware Images Classification: a CNN based Approach
- Focus: Querying Large Video Datasets with Low Latency and Low Cost
- Hierarchy-based Image Embeddings for Semantic Image Retrieval
- Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping
- Adversarial Examples for Semantic Segmentation and Object Detection
- Efficient Facial Representations for Age, Gender and Identity Recognition in Organizing Photo Albums using Multi-output CNN
- Domain Adaptive Transfer Learning with Specialist Models
- Forecasting with time series imaging
- Exploiting Local Features from Deep Networks for Image Retrieval
- Revisiting the Importance of Individual Units in CNNs via Ablation
- Neural Activation Constellations: Unsupervised Part Model Discovery with Convolutional Networks
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- VBPR: Visual Bayesian Personalized Ranking from Implicit Feedback
- Encoding High Dimensional Local Features by Sparse Coding Based Fisher Vectors
- Evaluating color texture descriptors under large variations of controlled lighting conditions
- Knowledge Distillation in Generations: More Tolerant Teachers Educate Better Students
- Webly Supervised Learning of Convolutional Networks
- Bilinear CNNs for Fine-grained Visual Recognition
- Automatic Hierarchical Classification of Kelps using Deep Residual Features
- A Probabilistic Quality Representation Approach to Deep Blind Image Quality Prediction
- Metric Learning with Adaptive Density Discrimination
- Phoneme-to-viseme mappings: the good, the bad, and the ugly
- DeepLogo: Hitting Logo Recognition with the Deep Neural Network Hammer
- Overcoming Small Minirhizotron Datasets Using Transfer Learning
- Label Efficient Learning of Transferable Representations across Domains and Tasks
- Multi-label Image Recognition by Recurrently Discovering Attentional Regions
- Cross-Domain Image Matching with Deep Feature Maps
- Learning From Noisy Large-Scale Datasets With Minimal Supervision
- Deep convolutional filter banks for texture recognition and segmentation
- Stacked Generative Adversarial Networks
- Convolutional Neural Networks: Ensemble Modeling, Fine-Tuning and Unsupervised Semantic Localization for Intraoperative CLE Images
- Unsupervised Semantic-based Aggregation of Deep Convolutional Features
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Leveraging Deep Graph-Based Text Representation for Sentiment Polarity Applications
- Learning Visual Clothing Style with Heterogeneous Dyadic Co-occurrences
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Heated-Up Softmax Embedding
- Part Detector Discovery in Deep Convolutional Neural Networks
- Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
- In-domain representation learning for remote sensing
- Temporal Relational Reasoning in Videos
- Yum-me: A Personalized Nutrient-based Meal Recommender System
- Fisher Kernel for Deep Neural Activations
- Deep Triplet Ranking Networks for One-Shot Recognition
- Pilot Comparative Study of Different Deep Features for Palmprint Identification in Low-Quality Images
- Maximum A Posteriori Estimation of Distances Between Deep Features in Still-to-Video Face Recognition
- VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization
- SIFT Meets CNN: A Decade Survey of Instance Retrieval
- Deep Convolutional Features for Image Based Retrieval and Scene Categorization
- Discrete and continuous representations and processing in deep learning: Looking forward
- Automatic Discovery and Optimization of Parts for Image Classification
- A Good Practice Towards Top Performance of Face Recognition: Transferred Deep Feature Fusion
- Preferences Prediction using a Gallery of Mobile Device based on Scene Recognition and Object Detection
- Smooth Neighbors on Teacher Graphs for Semi-supervised Learning
- Maximum-Margin Structured Learning with Deep Networks for 3D Human Pose Estimation
- Disentangling Image Distortions in Deep Feature Space
- What's Mine is Yours: Pretrained CNNs for Limited Training Sonar ATR
- Evaluating Contrastive Models for Instance-based Image Retrieval
- Learning with hidden variables
- Fully-Convolutional Siamese Networks for Object Tracking
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- A Deep Four-Stream Siamese Convolutional Neural Network with Joint Verification and Identification Loss for Person Re-detection
- Adversarial Reprogramming of Neural Networks
- Automatic localization and decoding of honeybee markers using deep convolutional neural networks
- Multimodal Classification for Analysing Social Media
- Learning Dense Convolutional Embeddings for Semantic Segmentation
- Image Credibility Analysis with Effective Domain Transferred Deep Networks
- Visualizing and Understanding Deep Texture Representations
- Learning Robust Bed Making using Deep Imitation Learning with DART
- Where to Focus: Query Adaptive Matching for Instance Retrieval Using Convolutional Feature Maps
- Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning
- Transfer Learning for Sequence Labeling Using Source Model and Target Data
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Toward Goal-Driven Neural Network Models for the Rodent Whisker-Trigeminal System
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Crime Mapping from Satellite Imagery via Deep Learning
- Transferring Autonomous Driving Knowledge on Simulated and Real Intersections
- Improving Image Clustering With Multiple Pretrained CNN Feature Extractors
- Transfer Learning for Material Classification using Convolutional Networks
- Stock Chart Pattern recognition with Deep Learning
- Explaining Explanations to Society
- Learning the Localization Function: Machine Learning Approach to Fingerprinting Localization
- Digging Deep into the layers of CNNs: In Search of How CNNs Achieve View Invariance
- Determinantal Point Processes for Mini-Batch Diversification
- RIGA: Covert and Robust White-Box Watermarking of Deep Neural Networks
- Local Feature Detectors, Descriptors, and Image Representations: A Survey
- Utilizing Deep Learning Towards Multi-modal Bio-sensing and Vision-based Affective Computing
- Personalized Driver Stress Detection with Multi-task Neural Networks using Physiological Signals
- A Framework for Searching for General Artificial Intelligence
- A Practical Guide to CNNs and Fisher Vectors for Image Instance Retrieval
- Deep Representation of Facial Geometric and Photometric Attributes for Automatic 3D Facial Expression Recognition
- Scalable Deep Learning Logo Detection
- Situational Object Boundary Detection
- From Image-level to Pixel-level Labeling with Convolutional Networks
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Estimated Depth Map Helps Image Classification
- Group Invariant Deep Representations for Image Instance Retrieval
- Unsupervised Category Discovery via Looped Deep Pseudo-Task Optimization Using a Large Scale Radiology Image Database
- Class Rectification Hard Mining for Imbalanced Deep Learning
- Pointwise Convolutional Neural Networks
- A Way out of the Odyssey: Analyzing and Combining Recent Insights for LSTMs
- Image search using multilingual texts: a cross-modal learning approach between image and text
- Toward Automatic Threat Recognition for Airport X-ray Baggage Screening with Deep Convolutional Object Detection
- Learning Visual Features from Large Weakly Supervised Data
- SA-CNN: Dynamic Scene Classification using Convolutional Neural Networks
- SVM and ELM: Who Wins? Object Recognition with Deep Convolutional Features from ImageNet
- Deep Motion Features for Visual Tracking
- What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?
- Computer-Aided Colorectal Tumor Classification in NBI Endoscopy Using CNN Features
- The More, the Better? A Study on Collaborative Machine Learning for DGA Detection
- Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Simultaneous Feature Aggregating and Hashing for Large-scale Image Search
- Is Image Super-resolution Helpful for Other Vision Tasks?
- Accelerating Deep Neural Networks with Spatial Bottleneck Modules
- Multi-modal Approach for Affective Computing
- Transductive Multi-label Zero-shot Learning
- Matching neural paths: transfer from recognition to correspondence search
- Unsupervised Feature Learning for Writer Identification and Writer Retrieval
- Unsupervised Joint Mining of Deep Features and Image Labels for Large-scale Radiology Image Categorization and Scene Recognition
- CuratorNet: Visually-aware Recommendation of Art Images
- Deep Epitomic Convolutional Neural Networks
- Conditional Deep Learning for Energy-Efficient and Enhanced Pattern Recognition
- An Out-of-the-box Full-network Embedding for Convolutional Neural Networks
- On Classification of Distorted Images with Deep Convolutional Neural Networks
- Semi-Supervising Learning, Transfer Learning, and Knowledge Distillation with SimCLR
- End-to-end Image Captioning Exploits Multimodal Distributional Similarity
- Bottom-Up and Top-Down Reasoning with Hierarchical Rectified Gaussians
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- Learning From Less Data: Diversified Subset Selection and Active Learning in Image Classification Tasks
- The Best of Both Worlds: Combining Data-independent and Data-driven Approaches for Action Recognition
- Cross-Domain Face Verification: Matching ID Document and Self-Portrait Photographs
- Complete the Look: Scene-based Complementary Product Recommendation
- Co-Regularized Deep Representations for Video Summarization
- Ad-Net: Audio-Visual Convolutional Neural Network for Advertisement Detection In Videos
- HEp-2 Cell Image Classification with Deep Convolutional Neural Networks
- Exploit Bounding Box Annotations for Multi-label Object Recognition
- Multi-Task Curriculum Transfer Deep Learning of Clothing Attributes
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Part-Stacked CNN for Fine-Grained Visual Categorization
- Target Driven Visual Navigation with Hybrid Asynchronous Universal Successor Representations
- Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents
- Visual Script and Language Identification
- Building Graph Representations of Deep Vector Embeddings
- Visual Fashion-Product Search at SK Planet
- Vision based body gesture meta features for Affective Computing
- Cost-effective Interactive Attention Learning with Neural Attention Processes
- Iterative Manifold Embedding Layer Learned by Incomplete Data for Large-scale Image Retrieval
- Learning Discriminative 3D Shape Representations by View Discerning Networks
- Kernelized Deep Convolutional Neural Network for Describing Complex Images
- Learning Image Conditioned Label Space for Multilabel Classification
- Towards an efficient deep learning model for musical onset detection
- Learning Image-Conditioned Dynamics Models for Control of Under-actuated Legged Millirobots
- A Deep Multi-task Learning Approach to Skin Lesion Classification
- On Catastrophic Interference in Atari 2600 Games
- Security Vulnerability Detection Using Deep Learning Natural Language Processing
- Semantic Hierarchical Priors for Intrinsic Image Decomposition
- Learning Manipulation under Physics Constraints with Visual Perception
- Exploiting Unsupervised Pre-training and Automated Feature Engineering for Low-resource Hate Speech Detection in Polish
- Object-Level Context Modeling For Scene Classification with Context-CNN
- Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-Shot In-Field Reflectance
- Context Augmentation for Convolutional Neural Networks
- Detection and Tracking of General Movable Objects in Large 3D Maps
- Towards CNN Map Compression for camera relocalisation
- How Useful is Region-based Classification of Remote Sensing Images in a Deep Learning Framework?
- MetaPerturb: Transferable Regularizer for Heterogeneous Tasks and Architectures
- Loop Closure Detection with RGB-D Feature Pyramid Siamese Networks
- Hierarchical Transfer Convolutional Neural Networks for Image Classification
- One-Shot Segmentation in Clutter
- Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions
- Periocular Recognition Using CNN Features Off-the-Shelf
- Class Subset Selection for Transfer Learning using Submodularity
- Supervised dimensionality reduction by a Linear Discriminant Analysis on pre-trained CNN features
- Learning to Select Pre-Trained Deep Representations with Bayesian Evidence Framework
- Attention Monitoring and Hazard Assessment with Bio-Sensing and Vision: Empirical Analysis Utilizing CNNs on the KITTI Dataset
- Recognition of Activities from Eye Gaze and Egocentric Video
- Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification
- Federated Few-Shot Learning with Adversarial Learning
- Blurred Images Lead to Bad Local Minima
- Query-free Clothing Retrieval via Implicit Relevance Feedback
- Compositional Model based Fisher Vector Coding for Image Classification
- Dynamic texture and scene classification by transferring deep image features
- Deep Collaborative Learning for Visual Recognition
- Adversarial Soft-detection-based Aggregation Network for Image Retrieval
- Counting Grid Aggregation for Event Retrieval and Recognition
- Multi-scale recognition with DAG-CNNs
- Retrieval of Family Members Using Siamese Neural Network
- LOTS about Attacking Deep Features
- Domain Adaptation with L2 constraints for classifying images from different endoscope systems
- Transfer Learning Based on AdaBoost for Feature Selection from Multiple ConvNet Layer Features
- Toward Multimodal Modeling of Emotional Expressiveness
- Modelling Temporal Information Using Discrete Fourier Transform for Recognizing Emotions in User-generated Videos
- Joint Learning of Distributed Representations for Images and Texts
- A Useful Motif for Flexible Task Learning in an Embodied Two-Dimensional Visual Environment
- An Analysis of Human-centered Geolocation
- Learning Neural Networks on SVD Boosted Latent Spaces for Semantic Classification
- Vision Recognition using Discriminant Sparse Optimization Learning
- Scalable Object Detection for Stylized Objects
- Learning a Complete Image Indexing Pipeline
- Multi-Scale Gradual Integration CNN for False Positive Reduction in Pulmonary Nodule Detection
- Weakly-Supervised Spatial Context Networks
- Compact Deep Aggregation for Set Retrieval
- Towards Efficient and Secure Delivery of Data for Deep Learning with Privacy-Preserving
- Towards Visual Feature Translation
- Location Recognition Over Large Time Lags
- Human-Centered Tools for Coping with Imperfect Algorithms during Medical Decision-Making
- Evaluating the Representational Hub of Language and Vision Models