Show and Tell: A Neural Image Caption Generator
arXiv:1411.4555
Abstract
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image. The model is trained to maximize the likelihood of the target description sentence given the training image. Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions. Our model is often quite accurate, which we verify both qualitatively and quantitatively. For instance, while the current state-of-the-art BLEU-1 score (the higher the better) on the Pascal dataset is 25, our approach yields 59, to be compared to human performance around 69. We also show BLEU-1 score improvements on Flickr30k, from 56 to 66, and on SBU, from 19 to 28. Lastly, on the newly released COCO dataset, we achieve a BLEU-4 of 27.7, which is the current state-of-the-art.
Cited by in corpus (44)
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Fully Connected Deep Structured Networks
- Neural Paraphrase Generation with Stacked Residual LSTM Networks
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- Long Short-Term Memory Over Tree Structures
- See, Hear, and Read: Deep Aligned Representations
- Person Search with Natural Language Description
- CIDEr: Consensus-based Image Description Evaluation
- Transferring Knowledge from a RNN to a DNN
- A Dataset for Movie Description
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Dual Attention Networks for Multimodal Reasoning and Matching
- Attribute Recognition by Joint Recurrent Learning of Context and Correlation
- Deep Recurrent Neural Networks for Acoustic Modelling
- Simple Image Description Generator via a Linear Phrase-Based Approach
- VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- GuessWhat?! Visual object discovery through multi-modal dialogue
- Semantic Compositional Networks for Visual Captioning
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Deep Learning Methods for Efficient Large Scale Video Labeling
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Recurrent Filter Learning for Visual Tracking
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Large-scale Video Classification guided by Batch Normalized LSTM Translator
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- Sketch2code: Generating a website from a paper mockup
- DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
- Learning to Disambiguate by Asking Discriminative Questions
- Examining Cooperation in Visual Dialog Models
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Few-Shot Object Recognition from Machine-Labeled Web Images
- SPICE: Semantic Propositional Image Caption Evaluation
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Intrinsic Image Captioning Evaluation
- Composing Text and Image for Image Retrieval - An Empirical Odyssey
- Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling
- A Sequential Set Generation Method for Predicting Set-Valued Outputs