154 citations · 315 across the 30 of their papers we have counts for
5 papers · 1 filter
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
Voice Conversion-based Privacy through Adversarial Information Hiding
Jacob J Webber, Oliver Watts, Gustav Eje Henter +2
Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper pre…
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
Yicheng Gu, Pablo Pérez Zarazaga, Chaoren Wang +4
Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant s…
Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks
Cassia Valentini-Botinhao, Manuel Sam Ribeiro, Oliver Watts +2
Automatically predicting the outcome of subjective listening tests is a challenging task. Ratings may vary from person to person even if preferences are consistent across listeners…
Transformation of low-quality device-recorded speech to high-quality speech using improved SEGAN model
Seyyed Saeed Sarfjoo, Xin Wang, Gustav Eje Henter +3
Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this res…