5 papers
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ulgen +4
Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted,…
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
Aurosweta Mahapatra, Ismail Rasim Ulgen, Berrak Sisman
Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to huma…
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
Ismail Rasim Ulgen, Zongyang Du, Junchen Lu +2
Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and wea…
Can Emotion Fool Anti-spoofing?
Aurosweta Mahapatra, Ismail Rasim Ulgen, Abinay Reddy Naini +2
Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness…
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
Philip H. Lee, Ismail Rasim Ulgen, Berrak Sisman
Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling t…