activity
20162022
most citedFeature learning for efficient ASR-free keyword spotting in low-resource languages

17 citations · 46 across the 22 of their papers we have counts for

collaborators

46 papers

eess.AS2022

TransFusion: Transcribing Speech with Multinomial Diffusion

Matthew Baas, Kevin Eloff, Herman Kamper

Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional t…

cs.CL2022

Towards visually prompted keyword localisation for zero-resource spoken languages

Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this ta…

cs.CL2022

YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding

Kayode Olaleye, Dan Oneata, Herman Kamper

Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is imposs…

cs.SD2022

GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from Diffusion Models

Matthew Baas, Herman Kamper

We propose AudioStyleGAN (ASGAN), a new generative adversarial network (GAN) for unconditional speech synthesis. As in the StyleGAN family of image synthesis models, ASGAN maps sam…

eess.AS202117 cited

Feature learning for efficient ASR-free keyword spotting in low-resource languages

Ewald van der Westhuizen, Herman Kamper, Raghav Menon +2

We consider feature learning for efficient keyword spotting that can be applied in severely under-resourced settings. The objective is to support humanitarian relief programmes by…

eess.AS20211 cited

Analyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing

Benjamin van Niekerk, Leanne Nortje, Matthew Baas +1

Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that line…