4 papers
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Meiying Melissa Chen, Anastasia Kuznetsova, Zhenyu Wang +1
Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS…
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
Anastasia Kuznetsova, Inseon Jang, Wootaek Lim +1
Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perceptio…
Discrete Audio Tokens: More Than a Survey!
Pooneh Mousavi, Gallil Maimon, Adel Moumen +18
Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and infere…
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha +3
This paper presents a new challenge that calls for zero-shot text-to-speech (TTS) systems to augment speech data for the downstream task, personalized speech enhancement (PSE), as…