activity
20242026
collaborators

16 papers

eess.AS2026

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

Sihang Nie, Jinxin Ji, Xiaofen Xing +4

While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms…

eess.AS2026

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Sihang Nie, Xiaofen Xing, Rui Xing +5

Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges t…

cs.AI2026

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2

Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generatio…

cs.SD2026

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

Jialong Mai, Jinxin Ji, Xiaofen Xing +2

Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus…

cs.SD2026

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

Jialong Mai, Xiaofen Xing, Xiangmin Xu

Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate co…

eess.AS2026

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS

Sihang Nie, Xiaofen Xing, Jingyuan Xing +2

Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging…