Imagination improves Multimodal Translation
arXiv:1705.04350
Abstract
We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.
Clarified main contributions, minor correction to Equation 8, additional comparisons in Table 2, added more related work
References in corpus (7)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Fast Domain Adaptation for Neural Machine Translation
- Incorporating Global Visual Features into Attention-Based Neural Machine Translation
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- Neural Machine Translation with Latent Semantic of Image and Text
- Keystroke dynamics as signal for shallow syntactic parsing
- Learning language through pictures
Cited by in corpus (18)
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Emergent Translation in Multi-Agent Communication
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning
- COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval
- Modulating and attending the source image during encoding improves Multimodal Translation
- Emergent Communication Pretraining for Few-Shot Machine Translation
- Visually Grounded Word Embeddings and Richer Visual Features for Improving Multimodal Neural Machine Translation
- Towards Multimodal Simultaneous Neural Machine Translation
- A Survey on Low-Resource Neural Machine Translation
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
- Generative Imagination Elevates Machine Translation
- Lessons learned in multilingual grounded language learning
- Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations
- Incorporating Chinese Radicals Into Neural Machine Translation: Deeper Than Character Level
- The MeMAD Submission to the WMT18 Multimodal Translation Task
- Probing Representations Learned by Multimodal Recurrent and Transformer Models