Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
arXiv:1504.06692
Abstract
In this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions. Using linguistic context and visual features, our method is able to efficiently hypothesize the semantic meaning of new words and add them to its word dictionary so that they can be used to describe images which contain these novel concepts. Our method has an image captioning module based on m-RNN with several improvements. In particular, we propose a transposed weight sharing scheme, which not only improves performance on image captioning, but also makes the model more suitable for the novel concept learning task. We propose methods to prevent overfitting the new concepts. In addition, three novel concept datasets are constructed for this new task. In the experiments, we show that our method effectively learns novel visual concepts from a few examples without disturbing the previously learned concepts. The project page is http://www.stat.ucla.edu/~junhua.mao/projects/child_learning.html
ICCV 2015 camera ready version. We add much more novel visual concepts in the NVC dataset and have released it, see http://www.stat.ucla.edu/~junhua.mao/projects/child_learning.html
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- ADADELTA: An Adaptive Learning Rate Method
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Explain Images with Multimodal Recurrent Neural Networks
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Learning Longer Memory in Recurrent Neural Networks
- Learning a Recurrent Visual Representation for Image Caption Generation
- Exploring Nearest Neighbor Approaches for Image Captioning
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- CIDEr: Consensus-based Image Description Evaluation
- Simple Image Description Generator via a Linear Phrase-Based Approach
- A Pooling Approach to Modelling Spatial Relations for Image Retrieval and Annotation
Cited by in corpus (26)
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Image Captioning with Semantic Attention
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Generating Visual Explanations
- Decoupled Novel Object Captioner
- Seeing with Humans: Gaze-Assisted Neural Image Captioning
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- Cascaded Revision Network for Novel Object Captioning
- Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects
- Interactive Grounded Language Acquisition and Generalization in a 2D World
- Intrinsic Relationship Reasoning for Small Object Detection
- SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
- Deep Nets: What have they ever done for Vision?
- Few-Shot Image Recognition by Predicting Parameters from Activations
- Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data
- MAT: A Multimodal Attentive Translator for Image Captioning
- A Semi-supervised Framework for Image Captioning
- A Comprehensive Survey of Deep Learning for Image Captioning
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- Title Generation for User Generated Videos
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval
- "Factual" or "Emotional": Stylized Image Captioning with Adaptive Learning and Attention
- Text-guided Attention Model for Image Captioning
- Transfer learning from language models to image caption generators: Better models may not transfer better
- Pointing Novel Objects in Image Captioning