4 papers
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
Yixuan Xiao, Florian Lux, Alejandro Pérez-González-de-Martos +1
Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulat…
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
Paul Mayer, Florian Lux, Alejandro Pérez-González-de-Martos +4
While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is par…
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
Xin Wang, Héctor Delgado, Hemlata Tak +26
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce…
High-Resolution Speech Restoration with Latent Diffusion Model
Tushar Dhyani, Florian Lux, Michele Mancusi +3
Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions fre…