2 papers
cs.CL2022
Modelling word learning and recognition using visually grounded speech
Danny Merkx, Sebastiaan Scholten, Stefan L. Frank +2
Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech f…
cs.CL2020
Learning to Recognise Words using Visually Grounded Speech
Sebastiaan Scholten, Danny Merkx, Odette Scharenborg
We investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which…