17 citations · 24 across the 9 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022
Multilingual Multimodal Learning with Machine Translated Text
Chen Qiu, Dan Oneata, Emanuele Bugliarello +2
Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL)…
cs.CL2022
YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
Kayode Olaleye, Dan Oneata, Herman Kamper
Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is imposs…