6 citations · 9 across the 13 of their papers we have counts for
6 papers · 1 filter
Translating speech with just images
Dan Oneata, Herman Kamper
Visually grounded speech models link speech to images. We extend this connection by linking images to text via an existing image captioning system, and as a result gain the ability…
Revisiting speech segmentation and lexicon learning with better features
Herman Kamper, Benjamin van Niekerk
We revisit a self-supervised method that segments unlabelled speech into word-like segments. We start from the two-stage duration-penalised dynamic programming method that performs…
Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices
Matthew Baas, Herman Kamper
Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output.…
Rhythm Modeling for Voice Conversion
Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in…
Voice Conversion With Just Nearest Neighbors
Matthew Baas, Benjamin van Niekerk, Herman Kamper
Any-to-any voice conversion aims to transform source speech into a target voice with just a few examples of the target speaker as a reference. Recent methods produce convincing con…
A Temporal Extension of Latent Dirichlet Allocation for Unsupervised Acoustic Unit Discovery
Werner van der Merwe, Herman Kamper, Johan du Preez
Latent Dirichlet allocation (LDA) is widely used for unsupervised topic modelling on sets of documents. No temporal information is used in the model. However, there is often a rela…