17 citations · 36 across the 20 of their papers we have counts for
Showing 2022Show all
3 papers · 1 filter
cs.CL2022
Multilingual Multimodal Learning with Machine Translated Text
Chen Qiu, Dan Oneata, Emanuele Bugliarello +2
Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL)…
cs.CL2022
YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
Kayode Olaleye, Dan Oneata, Herman Kamper
Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is imposs…
cs.SD2022
Improving Multimodal Speech Recognition by Data Augmentation and Speech Representations
Dan Oneata, Horia Cucu
Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated t…