CNN: Single-label to Multi-label
arXiv:1406.5726 · doi:10.1109/TPAMI.2015.2491929
Abstract
Convolutional Neural Network (CNN) has demonstrated promising performance in single-label image classification tasks. However, how CNN best copes with multi-label images still remains an open problem, mainly due to the complex underlying object layouts and insufficient multi-label training images. In this work, we propose a flexible deep CNN infrastructure, called Hypotheses-CNN-Pooling (HCP), where an arbitrary number of object segment hypotheses are taken as the inputs, then a shared CNN is connected with each hypothesis, and finally the CNN output results from different hypotheses are aggregated with max pooling to produce the ultimate multi-label predictions. Some unique characteristics of this flexible deep CNN infrastructure include: 1) no ground truth bounding box information is required for training; 2) the whole HCP infrastructure is robust to possibly noisy and/or redundant hypotheses; 3) no explicit hypothesis label is required; 4) the shared CNN may be well pre-trained with a large-scale single-label image dataset, e.g. ImageNet; and 5) it may naturally output multi-label prediction results. Experimental results on Pascal VOC2007 and VOC2012 multi-label image datasets well demonstrate the superiority of the proposed HCP infrastructure over other state-of-the-arts. In particular, the mAP reaches 84.2% by HCP only and 90.3% after the fusion with our complementary result in [47] based on hand-crafted features on the VOC2012 dataset, which significantly outperforms the state-of-the-arts with a large margin of more than 7%.
13 pages, 10 figures, 3 tables
References in corpus (3)
Cited by in corpus (108)
- STC: A Simple to Complex Framework for Weakly-supervised Semantic Segmentation
- Deep Label Distribution Learning with Label Ambiguity
- Diversified Visual Attention Networks for Fine-Grained Object Classification
- Multi-task CNN Model for Attribute Prediction
- CNN-RNN: A Unified Framework for Multi-label Image Classification
- Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
- Deep Ranking for Person Re-identification via Joint Representation Learning
- Learning to Discover Multi-Class Attentional Regions for Multi-Label Image Recognition
- Instance-Aware Hashing for Multi-Label Image Retrieval
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised Detection
- Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and Benchmark
- Light Field Salient Object Detection: A Review and Benchmark
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification
- Multi-Label Image Recognition with Graph Convolutional Networks
- Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
- An End-to-End Breast Tumour Classification Model Using Context-Based Patch Modelling- A BiLSTM Approach for Image Classification
- Learning Deep Latent Spaces for Multi-Label Classification
- Deep convolutional filter banks for texture recognition and segmentation
- Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
- Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach
- Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation
- Learning to Exploit the Prior Network Knowledge for Weakly-Supervised Semantic Segmentation
- HD-CNN: Hierarchical Deep Convolutional Neural Network for Large Scale Visual Recognition
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Deep FisherNet for Object Classification
- Tips, guidelines and tools for managing multi-label datasets: the mldr.datasets R package and the Cometa data repository
- What value do explicit high level concepts have in vision to language problems?
- Object Detection Networks on Convolutional Feature Maps
- Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
- Fisher Kernel for Deep Neural Activations
- Tensor Normalization and Full Distribution Training
- UntrimmedNets for Weakly Supervised Action Recognition and Detection
- Multi-layered Semantic Representation Network for Multi-label Image Classification
- Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition
- Joint Attention in Driver-Pedestrian Interaction: from Theory to Practice
- Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
- Exploring Categorical Regularization for Domain Adaptive Object Detection
- Approximate Fisher Kernels of non-iid Image Models for Image Categorization
- Weighting Scheme for a Pairwise Multi-label Classifier Based on the Fuzzy Confusion Matrix
- Computational Baby Learning
- Scale-aware Pixel-wise Object Proposal Networks
- GM-MLIC: Graph Matching based Multi-Label Image Classification
- Multi-label Image Classification using Adaptive Graph Convolutional Networks: from a Single Domain to Multiple Domains
- General Multi-label Image Classification with Transformers
- Annotation Order Matters: Recurrent Image Annotator for Arbitrary Length Image Tagging
- Learning Category Correlations for Multi-label Image Recognition with Graph Networks
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification
- Multi-Label Image Classification with Regional Latent Semantic Dependencies
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Indexing of CNN Features for Large Scale Image Search
- Deep Semantic Dictionary Learning for Multi-label Image Classification
- Object Level Deep Feature Pooling for Compact Image Representation
- A Correction Method of a Binary Classifier Applied to Multi-label Pairwise Models
- Holistic Interstitial Lung Disease Detection using Deep Convolutional Neural Networks: Multi-label Learning and Unordered Pooling
- sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification
- Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Elucidating image-to-set prediction: An analysis of models, losses and datasets
- The Shallow End: Empowering Shallower Deep-Convolutional Networks through Auxiliary Outputs
- Data-Free Knowledge Amalgamation via Group-Stack Dual-GAN
- Global Meets Local: Effective Multi-Label Image Classification via Category-Aware Weak Supervision
- Exploit Bounding Box Annotations for Multi-label Object Recognition
- Multi-label Ranking: Mining Multi-label and Label Ranking Data
- Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
- HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
- Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition
- A Facial Feature Discovery Framework for Race Classification Using Deep Learning
- Coarse to Fine: Multi-label Image Classification with Global/Local Attention
- LNEMLC: Label Network Embeddings for Multi-Label Classification
- Semi-Supervised Active Learning for COVID-19 Lung Ultrasound Multi-symptom Classification
- The iMaterialist Fashion Attribute Dataset
- Deep Attributes from Context-Aware Regional Neural Codes
- LID 2020: The Learning from Imperfect Data Challenge Results
- Room Geometry Estimation from Room Impulse Responses using Convolutional Neural Networks
- Learning Image Conditioned Label Space for Multilabel Classification
- Boosted GAN with Semantically Interpretable Information for Image Inpainting
- Learning to Point and Count
- Towards an Unequivocal Representation of Actions
- Training Object Detectors from Few Weakly-Labeled and Many Unlabeled Images
- SiMaN: Sign-to-Magnitude Network Binarization
- Ensemble of Part Detectors for Simultaneous Classification and Localization
- Instance-Level Salient Object Segmentation
- Instance-Aware Graph Convolutional Network for Multi-Label Classification
- Mining Mid-level Visual Patterns with Deep CNN Activations
- Progressive Representation Adaptation for Weakly Supervised Object Localization
- Amalgamating Filtered Knowledge: Learning Task-customized Student from Multi-task Teachers
- Multi-Target Prediction: A Unifying View on Problems and Methods
- Multi-Label Zero-Shot Human Action Recognition via Joint Latent Ranking Embedding
- Goal-driven text descriptions for images
- Object Proposal Generation using Two-Stage Cascade SVMs
- Multi-scale discriminative Region Discovery for Weakly-Supervised Object Localization
- Towards Accurate and Compact Architectures via Neural Architecture Transformer
- Disentangling Neural Architectures and Weights: A Case Study in Supervised Classification
- Multi-Label Learning from Single Positive Labels
- Shakeout: A New Approach to Regularized Deep Neural Network Training
- TS2C: Tight Box Mining with Surrounding Segmentation Context for Weakly Supervised Object Detection
- Reconstruction Regularized Deep Metric Learning for Multi-label Image Classification
- Transformer-based Dual Relation Graph for Multi-label Image Recognition
- Prototype Matching Networks for Large-Scale Multi-label Genomic Sequence Classification
- Unifying Relational Sentence Generation and Retrieval for Medical Image Report Composition
- A Scalable Pipelined Dataflow Accelerator for Object Region Proposals on FPGA Platform
- CascadeML: An Automatic Neural Network Architecture Evolution and Training Algorithm for Multi-label Classification
- Multiple Instance Learning Convolutional Neural Networks for Object Recognition
- Localizing Interpretable Multi-scale informative Patches Derived from Media Classification Task
- Kill Two Birds with One Stone: Weakly-Supervised Neural Network for Image Annotation and Tag Refinement
- Inferring Restaurant Styles by Mining Crowd Sourced Photos from User-Review Websites