2 papers
cs.SD2026
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
Xiaosu Su, Zihan Sun, Peilei Jia +1
Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style…
cs.SD2026
Hello-Chat: Towards Realistic Social Audio Interactions
Yueran Hou, Peilei Jia, Zihan Sun +5
Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer fr…