collaborators

8 papers

cs.SD2026

Taming Audio VAEs via Target-KL Regularization

Prem Seetharaman, Rithesh Kumar

Latent diffusion models have emerged as the dominant paradigm for many generation tasks including audio generation such as text-to-audio, text-to-music and text-to-speech. A key co…

cs.SD2026

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

William Chen, Prem Seetharaman, Rithesh Kumar +4

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can h…

cs.SD2026

TAC: Timestamped Audio Captioning

Sonal Kumar, Prem Seetharaman, Ke Chen +8

Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…

cs.SD2025

PromptSep: Generative Audio Separation via Multimodal Prompting

Yutong Wen, Ke Chen, Prem Seetharaman +7

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based…

eess.AS2025

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar +2

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity…

eess.AS2025

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

Justin Lovelace, Rithesh Kumar, Jiaqi Su +3

While generative Text-to-Speech (TTS) systems leverage vast ``in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitati…