2 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
Transformer-based Cascaded Multimodal Speech Translation
Zixiu Wu, Ozan Caglayan, Julia Ive +2
This paper describes the cascaded multimodal speech translation systems developed by Imperial College London for the IWSLT 2019 evaluation campaign. The architecture consists of an…
Imperial College London Submission to VATEX Video Captioning Task
Ozan Caglayan, Zixiu Wu, Pranava Madhyastha +2
This paper describes the Imperial College London team's submission to the 2019' VATEX video captioning challenge, where we first explore two sequence-to-sequence models, namely a r…
Predicting Actions to Help Predict Translations
Zixiu Wu, Julia Ive, Josiah Wang +2
We address the task of text translation on the How2 dataset using a state of the art transformer-based multimodal approach. The question we ask ourselves is whether visual features…
VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions
Pranava Madhyastha, Josiah Wang, Lucia Specia
We address the task of evaluating image description generation systems. We propose a novel image-aware metric for this task: VIFIDEL. It estimates the faithfulness of a generated c…