7 papers
Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS
Runwu Shi, Yujin Wang, Hongjin Song +1
Classifier-free guidance (CFG) is widely used in flow-matching-based zero-shot text-to-speech (TTS), where generation is typically controlled by two conditions: the target text and…
SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling
Weiqiao Shan, Ruixiang Mao, Yuang Li +10
Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive. MoE upcycling mitigates this cost by converting pretrained…
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
Haitao Li, Chunxiang Jin, Chenglin Li +3
Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result,…
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
Chengyuan Ma, Jiawei Jin, Ruijie Xiong +3
We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive au…
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
Dawei Huang, Yongjie Lv, Ruijie Xiong +2
Speech Emotion Recognition (SER) systems often assume congruence between vocal emotion and lexical semantics. However, in real-world interactions, acoustic-semantic conflict is com…
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
Canxiang Yan, Chunxiang Jin, Dawei Huang +22
Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech languag…