3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CL2022
YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
Kayode Olaleye, Dan Oneata, Herman Kamper
Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is imposs…
cs.CL2021
Attention-Based Keyword Localisation in Speech using Visual Grounding
Kayode Olaleye, Herman Kamper
Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, pr…
cs.CL2020★ 3 cited
Towards localisation of keywords in speech using weak supervision
Kayode Olaleye, Benjamin van Niekerk, Herman Kamper
Developments in weakly supervised and self-supervised models could enable speech technology in low-resource settings where full transcriptions are not available. We consider whethe…