8 papers
Do Speech Emphasis Models Generalize across Languages and Emotions?
Megan Wei, Deepali Aneja, Jiaqi Su +3
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speec…
TAC: Timestamped Audio Captioning
Sonal Kumar, Prem Seetharaman, Ke Chen +8
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
Jackie Lin, Jiaqi Su, Nishit Anand +3
Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability a…
PromptSep: Generative Audio Separation via Multimodal Prompting
Yutong Wen, Ke Chen, Prem Seetharaman +7
Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based…
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar +2
Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity…
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
Justin Lovelace, Rithesh Kumar, Jiaqi Su +3
While generative Text-to-Speech (TTS) systems leverage vast ``in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitati…