Context-aware Captions from Context-agnostic Supervision
arXiv:1701.02870
Abstract
We introduce an inference technique to produce discriminative context-aware image captions (captions that describe differences between images or visual concepts) using only generic context-agnostic training data (captions that describe a concept or an image in isolation). For example, given images and captions of "siamese cat" and "tiger cat", we generate language that describes the "siamese cat" in a way that distinguishes it from "tiger cat". Our key novelty is that we show how to do joint inference over a language model that is context-agnostic and a listener which distinguishes closely-related concepts. We first apply our technique to a justification task, namely to describe why an image contains a particular fine-grained category as opposed to another closely-related category of the CUB-200-2011 dataset. We then study discriminative image captioning to generate language that uniquely refers to one of two semantically-similar images in the COCO dataset. Evaluations with discriminative ground truth for justification and human studies for discriminative image captioning reveal that our approach outperforms baseline generative and speaker-listener approaches for discrimination.
Accepted to CVPR 2017 (Spotlight)
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show and Tell: A Neural Image Caption Generator
- Exploring Nearest Neighbor Approaches for Image Captioning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- Generating Visual Explanations
- CIDEr: Consensus-based Image Description Evaluation
- Reasoning About Pragmatics with Neural Listeners and Speakers
Cited by in corpus (15)
- Dialog-based Interactive Image Retrieval
- Discriminability objective for training descriptive captions
- Equal But Not The Same: Understanding the Implicit Relationship Between Persuasive Images and Text
- Unified Pragmatic Models for Generating and Following Instructions
- Pragmatically Informative Text Generation
- Semantic Explanations of Predictions
- Understanding Convolutional Networks with APPLE : Automatic Patch Pattern Labeling for Explanation
- Multilevel Context Representation for Improving Object Recognition
- Not All Words are Equal: Video-specific Information Loss for Video Captioning
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- Grounding Visual Explanations
- Group-based Distinctive Image Captioning with Memory Attention
- Reference-Centric Models for Grounded Collaborative Dialogue
- Informative Object Annotations: Tell Me Something I Don't Know