Predicting Deep Zero-Shot Convolutional Neural Networks using Textual Descriptions
arXiv:1506.00511
Abstract
One of the main challenges in Zero-Shot Learning of visual categories is gathering semantic attributes to accompany images. Recent work has shown that learning from textual descriptions, such as Wikipedia articles, avoids the problem of having to explicitly define these attributes. We present a new model that can classify unseen categories from their textual description. Specifically, we use text features to predict the output weights of both the convolutional and the fully connected layers in a deep convolutional neural network (CNN). We take advantage of the architecture of CNNs and learn features at different layers, rather than just learning an embedding space for both modalities, as is common with existing approaches. The proposed model also allows us to automatically generate a list of pseudo- attributes for each visual category consisting of words from Wikipedia articles. We train our models end-to-end us- ing the Caltech-UCSD bird and flower datasets and evaluate both ROC and Precision-Recall curves. Our empirical results show that the proposed model significantly outperforms previous methods.
Correct the typos in table 1 regarding [5]. To appear in ICCV 2015
References in corpus (4)
Cited by in corpus (36)
- Delta-encoder: an effective sample synthesis method for few-shot object recognition
- Representation Learning for Natural Language Processing
- Preserving Semantic Relations for Zero-Shot Learning
- Towards Zero-shot Sign Language Recognition
- Zero-Shot Learning and its Applications from Autonomous Vehicles to COVID-19 Diagnosis: A Review
- Zero-Shot Learning by Generating Pseudo Feature Representations
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- Recent Advances in Zero-shot Recognition
- Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction
- Look, Listen and Learn
- Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs
- Semi-supervised Zero-Shot Learning by a Clustering-based Approach
- Zero-Shot Learning posed as a Missing Data Problem
- Online but Accurate Inference for Latent Variable Models with Local Gibbs Sampling
- Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning
- Fine-Grained Zero-Shot Learning with DNA as Side Information
- Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
- Multi-Cue Zero-Shot Learning with Strong Supervision
- Learning Joint Feature Adaptation for Zero-Shot Recognition
- An inner-loop free solution to inverse problems using deep neural networks
- Predicting Visual Exemplars of Unseen Classes for Zero-Shot Learning
- Large-Scale Visual Relationship Understanding
- Zero-Shot Learning via Joint Latent Similarity Embedding
- Combinets: Creativity via Recombination of Neural Networks
- Less is more: zero-shot learning from online textual documents with noise suppression
- MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
- Zero and Few Shot Learning with Semantic Feature Synthesis and Competitive Learning
- Zero-Shot Learning via Category-Specific Visual-Semantic Mapping
- Agent-Centric Risk Assessment: Accident Anticipation and Risky Region Localization
- Joint Dictionaries for Zero-Shot Learning
- Single-View 3D Object Reconstruction from Shape Priors in Memory
- Imaginative Walks: Generative Random Walk Deviation Loss for Improved Unseen Learning Representation
- Infinite-Label Learning with Semantic Output Codes
- Automatic Discovery, Association Estimation and Learning of Semantic Attributes for a Thousand Categories
- Analyzing First-Person Stories Based on Socializing, Eating and Sedentary Patterns
- Zero-Shot Learning with Sparse Attribute Propagation