Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
arXiv:1609.06647 · doi:10.1109/TPAMI.2016.2587640
Abstract
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image. The model is trained to maximize the likelihood of the target description sentence given the training image. Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions. Our model is often quite accurate, which we verify both qualitatively and quantitatively. Finally, given the recent surge of interest in this task, a competition was organized in 2015 using the newly released COCO dataset. We describe and analyze the various improvements we applied to our own baseline and show the resulting performance in the competition, which we won ex-aequo with a team from Microsoft Research, and provide an open source implementation in TensorFlow.
arXiv admin note: substantial text overlap with arXiv:1411.4555
References in corpus (9)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Recurrent Neural Network Regularization
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Explain Images with Multimodal Recurrent Neural Networks
- Exploring Nearest Neighbor Approaches for Image Captioning
- Technical Report: Image Captioning with Semantically Similar Images
Cited by in corpus (27)
- 3G structure for image caption generation
- Learning Semantic Concepts and Order for Image and Sentence Matching
- Trading Off Diversity and Quality in Natural Language Generation
- Position Focused Attention Network for Image-Text Matching
- Generating Diverse and Meaningful Captions
- Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
- TPsgtR: Neural-Symbolic Tensor Product Scene-Graph-Triplet Representation for Image Captioning
- Multi-modal Discriminative Model for Vision-and-Language Navigation
- Multimodal Memory Modelling for Video Captioning
- Self-Segregating and Coordinated-Segregating Transformer for Focused Deep Multi-Modular Network for Visual Question Answering
- Better Text Understanding Through Image-To-Text Transfer
- Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
- Query by Semantic Sketch
- SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
- SuperCaptioning: Image Captioning Using Two-dimensional Word Embedding
- Towards Unique and Informative Captioning of Images
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Improving Image Captioning by Leveraging Knowledge Graphs
- Modularized Textual Grounding for Counterfactual Resilience
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- Recognizing and Curating Photo Albums via Event-Specific Image Importance
- Event Recognition with Automatic Album Detection based on Sequential Processing, Neural Attention and Image Captioning
- Towards Embodied Scene Description
- Building Chatbots from Forum Data: Model Selection Using Question Answering Metrics
- Towards ECDSA key derivation from deep embeddings for novel Blockchain applications
- Better Understanding Hierarchical Visual Relationship for Image Caption