activity
20182022
collaborators

6 papers

cs.CL2022

Modelling word learning and recognition using visually grounded speech

Danny Merkx, Sebastiaan Scholten, Stefan L. Frank +2

Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech f…

cs.CL2022

Seeing the advantage: visually grounding word embeddings to better capture human semantic knowledge

Danny Merkx, Stefan L. Frank, Mirjam Ernestus

Distributional semantic models capture word-level meaning that is useful in many natural language processing tasks and have even been shown to capture cognitive aspects of word mea…

cs.CL2020

Learning to Recognise Words using Visually Grounded Speech

Sebastiaan Scholten, Danny Merkx, Odette Scharenborg

We investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which…

cs.CL2019

Language learning using Speech to Image retrieval

Danny Merkx, Stefan L. Frank, Mirjam Ernestus

Humans learn language by interaction with their environment and listening to other humans. It should also be possible for computational models to learn language directly from speec…

cs.CL2019

Learning semantic sentence representations from visually grounded language without lexical knowledge

Danny Merkx, Stefan Frank

Current approaches to learning semantic representations of sentences often use prior word-level knowledge. The current study aims to leverage visual information in order to capture…

cs.CL2018

Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

Odette Scharenborg, Laurent Besacier, Alan Black +16

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and word…