1 paper
Sebastiaan Scholten, Danny Merkx, Odette Scharenborg
We investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which…