activity
20162025
most citedFeature learning for efficient ASR-free keyword spotting in low-resource languages

17 citations · 57 across the 41 of their papers we have counts for

collaborators
Showing 2023Show all

9 papers · 1 filter

eess.AS2023

Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices

Matthew Baas, Herman Kamper

Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output.…

eess.AS2023

Rhythm Modeling for Voice Conversion

Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in…

eess.AS2023★ 1 cited

Leveraging multilingual transfer for unsupervised semantic acoustic word embeddings

Christiaan Jacobs, Herman Kamper

Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have si…

eess.AS2023★ 1 cited

Disentanglement in a GAN for Unconditional Speech Synthesis

Matthew Baas, Herman Kamper

Can we develop a model that can synthesize realistic speech directly from a latent space, without explicit conditioning? Despite several efforts over the last decade, previous adve…

eess.AS2023

Visually grounded few-shot word learning in low-resource settings

Leanne Nortje, Dan Oneata, Herman Kamper

We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken quer…

cs.CL2023★ 1 cited

Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili

Christiaan Jacobs, Nathanaël Carraz Rakotonirina, Everlyn Asiko Chimoto +2

We consider hate speech detection through keyword spotting on radio broadcasts. One approach is to build an automatic speech recognition (ASR) system for the target low-resource la…