17 citations · 57 across the 41 of their papers we have counts for
9 papers · 1 filter
Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices
Matthew Baas, Herman Kamper
Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output.…
Rhythm Modeling for Voice Conversion
Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in…
Leveraging multilingual transfer for unsupervised semantic acoustic word embeddings
Christiaan Jacobs, Herman Kamper
Acoustic word embeddings (AWEs) are fixed-dimensional vector representations of speech segments that encode phonetic content so that different realisations of the same word have si…
Disentanglement in a GAN for Unconditional Speech Synthesis
Matthew Baas, Herman Kamper
Can we develop a model that can synthesize realistic speech directly from a latent space, without explicit conditioning? Despite several efforts over the last decade, previous adve…
Visually grounded few-shot word learning in low-resource settings
Leanne Nortje, Dan Oneata, Herman Kamper
We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken quer…
Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili
Christiaan Jacobs, Nathanaël Carraz Rakotonirina, Everlyn Asiko Chimoto +2
We consider hate speech detection through keyword spotting on radio broadcasts. One approach is to build an automatic speech recognition (ASR) system for the target low-resource la…