7 citations · 14 across the 11 of their papers we have counts for
4 papers · 1 filter
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
Yixuan Xiao, Florian Lux, Alejandro Pérez-González-de-Martos +1
Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulat…
High-Resolution Speech Restoration with Latent Diffusion Model
Tushar Dhyani, Florian Lux, Michele Mancusi +3
Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions fre…
Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal Directions
Florian Lux, Pascal Tilli, Sarina Meyer +1
Customizing voice and speaking style in a speech synthesis system with intuitive and fine-grained controls is challenging, given that little data with appropriate labels is availab…
Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy
Sarina Meyer, Pascal Tilli, Pavel Denisov +3
In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes wit…