Multimodal Pivots for Image Caption Translation
arXiv:1601.03916 · doi:10.18653/v1/p16-1227
Abstract
We present an approach to improve statistical machine translation of image descriptions by multimodal pivots defined in visual space. The key idea is to perform image retrieval over a database of images that are captioned in the target language, and use the captions of the most similar images for crosslingual reranking of translation outputs. Our approach does not depend on the availability of large amounts of in-domain parallel data, but only relies on available large datasets of monolingually captioned images, and on state-of-the-art convolutional neural networks to compute image similarities. Our experimental evaluation shows improvements of 1 BLEU point over strong baselines.
Final version, accepted at ACL 2016. New section on Human Evaluation
References in corpus (1)
Cited by in corpus (22)
- WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
- Imagination improves Multimodal Translation
- Neural Extractive Summarization with Side Information
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Emergent Translation in Multi-Agent Communication
- Unpaired Image Captioning via Scene Graph Alignments
- One Sentence One Model for Neural Machine Translation
- Using Artificial Tokens to Control Languages for Multilingual Image Caption Generation
- Doubly Attentive Transformer Machine Translation
- Neural Machine Translation with Latent Semantic of Image and Text
- COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval
- Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task
- Towards Multimodal Simultaneous Neural Machine Translation
- Zero-Resource Neural Machine Translation with Multi-Agent Communication Game
- Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards
- Zero-resource Machine Translation by Multimodal Encoder-decoder Network with Multimedia Pivot
- Practical Comparable Data Collection for Low-Resource Languages via Images
- FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions
- Neural Machine Translation: A Review and Survey
- MULE: Multimodal Universal Language Embedding
- Read, Listen, and See: Leveraging Multimodal Information Helps Chinese Spell Checking
- Bootstrapping Disjoint Datasets for Multilingual Multimodal Representation Learning