3 papers
cs.SD2025
UniSep: Universal Target Audio Separation with Language Models at Scale
Yuanyuan Wang, Hangting Chen, Dongchao Yang +7
We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep…
cs.SD2025
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
Weidong Chen, Shan Yang, Guangzhi Li +1
Controlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current…
cs.SD2023
SECap: Speech Emotion Captioning with Large Language Model
Yaoxun Xu, Hangting Chen, Jianwei Yu +6
Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech e…