activity
20162026
most citedUnsupervised neural and Bayesian models for zero-resource speech processing

6 citations · 9 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build…

cs.CL2026

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

Stephen McIntosh, Reuben Smit, Daisuke Saito +2

L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. Moreover, assessm…

cs.CL2024

Visually Grounded Speech Models have a Mutual Exclusivity Bias

Leanne Nortje, Dan Oneaţă, Yevgen Matusevych +1

When children learn new words, they employ constraints such as the mutual exclusivity (ME) bias: a novel word is mapped to a novel object rather than a familiar one. This bias has…

cs.CL20231 cited

Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili

Christiaan Jacobs, Nathanaël Carraz Rakotonirina, Everlyn Asiko Chimoto +2

We consider hate speech detection through keyword spotting on radio broadcasts. One approach is to build an automatic speech recognition (ASR) system for the target low-resource la…

cs.CL2023

Visually grounded few-shot word acquisition with fewer shots

Leanne Nortje, Benjamin van Niekerk, Herman Kamper

We propose a visually grounded speech model that acquires new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken qu…

cs.CL2023

Mitigating Catastrophic Forgetting for Few-Shot Spoken Word Classification Through Meta-Learning

Ruan van der Merwe, Herman Kamper

We consider the problem of few-shot spoken word classification in a setting where a model is incrementally introduced to new word classes. This would occur in a user-defined keywor…