Exploring Nearest Neighbor Approaches for Image Captioning
arXiv:1505.04467
Abstract
We explore a variety of nearest neighbor baseline approaches for image captioning. These approaches find a set of nearest neighbor images in the training set from which a caption may be borrowed for the query image. We select a caption for the query image by finding the caption that best represents the "consensus" of the set of candidate captions gathered from the nearest neighbor images. When measured by automatic evaluation metrics on the MS COCO caption evaluation server, these approaches perform as well as many recent approaches that generate novel captions. However, human studies show that a method that generates novel captions is still preferred over the nearest neighbor approach.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Explain Images with Multimodal Recurrent Neural Networks
- CIDEr: Consensus-based Image Description Evaluation
- Simple Image Description Generator via a Linear Phrase-Based Approach
Cited by in corpus (8)
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- Text-guided Attention Model for Image Captioning
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- Tinkering Under the Hood: Interactive Zero-Shot Learning with Net Surgery
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- SPICE: Semantic Propositional Image Caption Evaluation