6 papers
Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference
Yaohan Guan, Yen-Ju Lu, Yuzhe Wang +5
Inferring speaker relationships from spoken conversations is an important step towards socially aware speech understanding. However, this task remains underexplored, and supervised…
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Mingrui Liang, Thomas Thebaud, Lukasz Wojciak +4
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing…
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
Tianyu Cao, Helin Wang, Ari Frummer +7
Recent advances in generative models, such as diffusion and flow matching, have shown strong performance in audio tasks. However, speech enhancement (SE) models are typically train…
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Helin Wang, Jiarui Hai, Dading Chong +11
Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to…
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
Helin Wang, Jiarui Hai, Dongchao Yang +7
Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary aud…
Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition
Piotr Żelasko, Siyuan Feng, Laureano Moro Velazquez +5
The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a writt…