Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
arXiv:1609.06647 · doi:10.1109/TPAMI.2016.2587640
Abstract
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image. The model is trained to maximize the likelihood of the target description sentence given the training image. Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions. Our model is often quite accurate, which we verify both qualitatively and quantitatively. Finally, given the recent surge of interest in this task, a competition was organized in 2015 using the newly released COCO dataset. We describe and analyze the various improvements we applied to our own baseline and show the resulting performance in the competition, which we won ex-aequo with a team from Microsoft Research, and provide an open source implementation in TensorFlow.
arXiv admin note: substantial text overlap with arXiv:1411.4555
References in corpus (9)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Recurrent Neural Network Regularization
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Explain Images with Multimodal Recurrent Neural Networks
- Exploring Nearest Neighbor Approaches for Image Captioning
- Technical Report: Image Captioning with Semantically Similar Images
Cited by in corpus (49)
- Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval
- Multimodal Deep Learning Framework for Image Popularity Prediction on Social Media
- Supervised Contrastive Learning for Multimodal Unreliable News Detection in COVID-19 Pandemic
- Comparative evaluation of CNN architectures for Image Caption Generation
- 3G structure for image caption generation
- Discrete and continuous representations and processing in deep learning: Looking forward
- Detecting Hateful Memes Using a Multimodal Deep Ensemble
- Learning Semantic Concepts and Order for Image and Sentence Matching
- Trading Off Diversity and Quality in Natural Language Generation
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
- Position Focused Attention Network for Image-Text Matching
- Generating Diverse and Meaningful Captions
- SIGAN: A Novel Image Generation Method for Solar Cell Defect Segmentation and Augmentation
- Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
- TPsgtR: Neural-Symbolic Tensor Product Scene-Graph-Triplet Representation for Image Captioning
- Multi-modal Discriminative Model for Vision-and-Language Navigation
- Dissecting the Meme Magic: Understanding Indicators of Virality in Image Memes
- Multimodal Memory Modelling for Video Captioning
- Self-Segregating and Coordinated-Segregating Transformer for Focused Deep Multi-Modular Network for Visual Question Answering
- Fast Sequence Generation with Multi-Agent Reinforcement Learning
- Better Text Understanding Through Image-To-Text Transfer
- Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
- Query by Semantic Sketch
- Towards Overcoming False Positives in Visual Relationship Detection
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image Captioning
- SuperCaptioning: Image Captioning Using Two-dimensional Word Embedding
- SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
- Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Towards Unique and Informative Captioning of Images
- SuperOCR: A Conversion from Optical Character Recognition to Image Captioning
- Modularized Textual Grounding for Counterfactual Resilience
- Improving Image Captioning by Leveraging Knowledge Graphs
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- Event Recognition with Automatic Album Detection based on Sequential Processing, Neural Attention and Image Captioning
- Towards Embodied Scene Description
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- There Once Was a Really Bad Poet, It Was Automated but You Didn't Know It
- Graphine: A Dataset for Graph-aware Terminology Definition Generation
- Recognizing and Curating Photo Albums via Event-Specific Image Importance
- Symbolic music generation conditioned on continuous-valued emotions
- Sentence Semantic Regression for Text Generation
- Better Understanding Hierarchical Visual Relationship for Image Caption
- Experimenting with Self-Supervision using Rotation Prediction for Image Captioning
- Integrating Deep Learning and Augmented Reality to Enhance Situational Awareness in Firefighting Environments
- Towards ECDSA key derivation from deep embeddings for novel Blockchain applications
- Caption Enriched Samples for Improving Hateful Memes Detection
- Characterization and recognition of handwritten digits using Julia
- Building Chatbots from Forum Data: Model Selection Using Question Answering Metrics