64 citations · 133 across the 9 of their papers we have counts for
8 papers · 1 filter
Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module
Adam Gabryś, Goeric Huybrechts, Manuel Sam Ribeiro +6
State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data,…
Cross-speaker style transfer for text-to-speech using data augmentation
Manuel Sam Ribeiro, Julian Roth, Giulia Comini +3
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive d…
Automatic audiovisual synchronisation for ultrasound tongue imaging
Aciel Eshky, Joanne Cleland, Manuel Sam Ribeiro +3
Ultrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therap…
Silent versus modal multi-speaker speech recognition from ultrasound and video
Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond +1
We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evalu…
Exploiting ultrasound tongue imaging for the automatic detection of speech articulation errors
Manuel Sam Ribeiro, Joanne Cleland, Aciel Eshky +2
Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical interv…
TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos
Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang +4
We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording s…