Discriminability objective for training descriptive captions
arXiv:1803.04376
Abstract
One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning.
CVPR2018
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Actor-Critic Sequence Training for Image Captioning
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- CIDEr: Consensus-based Image Description Evaluation
- Modeling Context in Referring Expressions
- Dual Attention Networks for Multimodal Reasoning and Matching
- Boosting Image Captioning with Attributes
- Teaching Machines to Describe Images via Natural Language Feedback
- GuessWhat?! Visual object discovery through multi-modal dialogue
- Comprehension-guided referring expressions
- Learning to Disambiguate by Asking Discriminative Questions
Cited by in corpus (35)
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- Show, Recall, and Tell: Image Captioning with Recall Mechanism
- Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
- Reasoning Visual Dialogs with Structural and Partial Observations
- Describing like humans: on diversity in image captioning
- On Hallucination and Predictive Uncertainty in Conditional Language Generation
- Multi-modal Ensemble Models for Predicting Video Memorability
- Cross-Modal Graph with Meta Concepts for Video Captioning
- Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Fine-Grained Image Captioning with Global-Local Discriminative Objective
- Towards Diverse and Accurate Image Captions via Reinforcing Determinantal Point Process
- Context and Attribute Grounded Dense Captioning
- Generating Question Relevant Captions to Aid Visual Question Answering
- Expressing Visual Relationships via Language
- Human-like Controllable Image Captioning with Verb-specific Semantic Roles
- Better Captioning with Sequence-Level Exploration
- Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation Learning
- Detection and Description of Change in Visual Streams
- Graph Edit Distance Reward: Learning to Edit Scene Graph
- Structure-Aware Generation Network for Recipe Generation from Images
- Diversifying Dialogue Generation with Non-Conversational Text
- Towards Unique and Informative Captioning of Images
- Structural and Functional Decomposition for Personality Image Captioning in a Communication Game
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- Group-based Distinctive Image Captioning with Memory Attention
- Intrinsic Image Captioning Evaluation
- Bayesian Attention Modules
- Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder
- Scene-based Factored Attention for Image Captioning
- End-to-End Learning Using Cycle Consistency for Image-to-Caption Transformations
- Informative Object Annotations: Tell Me Something I Don't Know
- Contrastive Semantic Similarity Learning for Image Captioning Evaluation with Intrinsic Auto-encoder
- Counterfactual Maximum Likelihood Estimation for Training Deep Networks