5 papers
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
Sihang Nie, Jinxin Ji, Xiaofen Xing +4
While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms…
HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
Sihang Nie, Xiaofen Xing, Rui Xing +5
Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges t…
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2
Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generatio…
NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
Jialong Mai, Jinxin Ji, Xiaofen Xing +2
Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus…
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
Zikai Huang, Yihan Zhou, Xuemiao Xu +4
Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emot…