Language Models for Image Captioning: The Quirks and What Works
arXiv:1505.01809
Abstract
Two recent approaches have achieved state-of-the-art results in image captioning. The first uses a pipelined process where a set of candidate words is generated by a convolutional neural network (CNN) trained on images, and then a maximum entropy (ME) language model is used to arrange these words into a coherent sentence. The second uses the penultimate activation layer of the CNN as input to a recurrent neural network (RNN) that then generates the caption sequence. In this paper, we compare the merits of these different language modeling approaches for the first time by using the same state-of-the-art CNN as input. We examine issues in the different approaches, including linguistic irregularities, caption repetition, and data set overlap. By combining key aspects of the ME and RNN methods, we achieve a new record performance over previously published results on the benchmark COCO dataset. However, the gains we see in BLEU do not translate to human judgments.
See http://research.microsoft.com/en-us/projects/image_captioning for project information
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Explain Images with Multimodal Recurrent Neural Networks
- Show and Tell: A Neural Image Caption Generator
- Learning a Recurrent Visual Representation for Image Caption Generation
- Deep Visual-Semantic Alignments for Generating Image Descriptions
Cited by in corpus (53)
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Exploring Nearest Neighbor Approaches for Image Captioning
- Multimodal Machine Learning: A Survey and Taxonomy
- AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
- Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
- Towards Diverse and Natural Image Descriptions via a Conditional GAN
- Survey on the attention based RNN model and its applications in computer vision
- Actor-Critic Sequence Training for Image Captioning
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Seeing with Humans: Gaze-Assisted Neural Image Captioning
- Stacked Cross Attention for Image-Text Matching
- Rich Image Captioning in the Wild
- Boosting Image Captioning with Attributes
- Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects
- What value do explicit high level concepts have in vision to language problems?
- Image Inspired Poetry Generation in XiaoIce
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- Partially-Supervised Image Captioning
- Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
- Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
- Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
- Semantic Compositional Networks for Visual Captioning
- Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions
- Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition
- MAT: A Multimodal Attentive Translator for Image Captioning
- A Comprehensive Survey of Deep Learning for Image Captioning
- Going Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- Reflective Decoding Network for Image Captioning
- End-to-end Image Captioning Exploits Multimodal Distributional Similarity
- Tensor Product Generation Networks for Deep NLP Modeling
- Fundamental principles of cortical computation: unsupervised learning with prediction, compression and feedback
- Phrase-based Image Captioning with Hierarchical LSTM Model
- Generating Diverse and Informative Natural Language Fashion Feedback
- Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioning
- Gaussian Smoothen Semantic Features (GSSF) -- Exploring the Linguistic Aspects of Visual Captioning in Indian Languages (Bengali) Using MSCOCO Framework
- Embodied Language Grounding with 3D Visual Feature Representations
- Intention Oriented Image Captions with Guiding Objects
- Geometry-Entangled Visual Semantic Transformer for Image Captioning
- Empirical Autopsy of Deep Video Captioning Frameworks
- Improving Classification by Improving Labelling: Introducing Probabilistic Multi-Label Object Interaction Recognition
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- SPICE: Semantic Propositional Image Caption Evaluation
- Middle-Out Decoding
- Towards Understanding End-of-trip Instructions in a Taxi Ride Scenario
- Multi-Glimpse Network: A Robust and Efficient Classification Architecture based on Recurrent Downsampled Attention
- A Weighted Multi-Criteria Decision Making Approach for Image Captioning
- Predicting Drug Responses by Propagating Interactions through Text-Enhanced Drug-Gene Networks
- Automated Image Captioning for Rapid Prototyping and Resource Constrained Environments
- The Long-Short Story of Movie Description