2 papers
cs.SD2026
Audio ControlNet for Fine-Grained Audio Generation and Editing
Haina Zhu, Yao Xiao, Xiquan Li +5
We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over at…
cs.SD2025
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
Pengchao Feng, Yao Xiao, Ziyang Ma +5
Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted t…