Does Multimodality Help Human and Machine for Translation and Image Captioning?
arXiv:1605.09186 · doi:10.18653/v1/W16-2358
Abstract
This paper presents the systems developed by LIUM and CVC for the WMT16 Multimodal Machine Translation challenge. We explored various comparative methods, namely phrase-based systems and attentional recurrent neural networks models trained using monomodal or multimodal data. We also performed a human evaluation in order to estimate the usefulness of multimodal data for human machine translation and image description generation. Our systems obtained the best results for both tasks according to the automatic evaluation metrics BLEU and METEOR.
7 pages, 2 figures, v4: Small clarification in section 4 title and content
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Distributed Representations of Sentences and Documents
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Multilingual Image Description with Neural Sequence Models
Cited by in corpus (8)
- An empirical study on the effectiveness of images in Multimodal Neural Machine Translation
- Zero-Resource Translation with Multi-Lingual Neural Machine Translation
- Search Engine Guided Non-Parametric Neural Machine Translation
- Multimodal Compact Bilinear Pooling for Multimodal Neural Machine Translation
- UMONS Submission for WMT18 Multimodal Translation Task
- Visually Grounded Word Embeddings and Richer Visual Features for Improving Multimodal Neural Machine Translation
- Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
- Supervised Visual Attention for Simultaneous Multimodal Machine Translation