activity
20172026
most citedAnalyzing Input and Output Representations for Speech-Driven Gesture Generation

154 citations · 315 across the 30 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction

Téo Guichoux, Théodor Lemerle, Shivam Mehta +5

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…

cs.SD2024

Voice Conversion-based Privacy through Adversarial Information Hiding

Jacob J Webber, Oliver Watts, Gustav Eje Henter +2

Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper pre…

cs.SD2024

HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters

Yicheng Gu, Pablo Pérez Zarazaga, Chaoren Wang +4

Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant s…

cs.SD20223 cited

Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks

Cassia Valentini-Botinhao, Manuel Sam Ribeiro, Oliver Watts +2

Automatically predicting the outcome of subjective listening tests is a challenging task. Ratings may vary from person to person even if preferences are consistent across listeners…

cs.SD2019

Transformation of low-quality device-recorded speech to high-quality speech using improved SEGAN model

Seyyed Saeed Sarfjoo, Xin Wang, Gustav Eje Henter +3

Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this res…