16 papers
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
Sihang Nie, Jinxin Ji, Xiaofen Xing +4
While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms…
HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
Sihang Nie, Xiaofen Xing, Rui Xing +5
Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges t…
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2
Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generatio…
NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
Jialong Mai, Jinxin Ji, Xiaofen Xing +2
Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus…
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control
Jialong Mai, Xiaofen Xing, Xiangmin Xu
Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate co…
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
Sihang Nie, Xiaofen Xing, Jingyuan Xing +2
Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging…