8 papers
Do Speech Emphasis Models Generalize across Languages and Emotions?
Megan Wei, Deepali Aneja, Jiaqi Su +3
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speec…
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
Jackie Lin, Jiaqi Su, Nishit Anand +3
Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability a…
PromptSep: Generative Audio Separation via Multimodal Prompting
Yutong Wen, Ke Chen, Prem Seetharaman +7
Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based…
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar +2
Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity…
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
Justin Lovelace, Rithesh Kumar, Jiaqi Su +3
While generative Text-to-Speech (TTS) systems leverage vast ``in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitati…
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
Patrick O'Reilly, Zeyu Jin, Jiaqi Su +1
In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is…