14 papers
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
Pengchao Feng, Chao-Hong Tan, Qian Chen +3
Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoke…
BareWave: Waveform-Native Flow-Matching Text-to-Speech
Wei Fan, Chao-Hong Tan, Qian Chen +5
Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality syst…
UniVocal: Unified Speech-Singing Code-Switching Synthesis
Yufei Shi, Qian Chen, Wen Wang +3
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
Yifu Chen, Shengpeng Ji, Qian Chen +9
End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems.…
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
Yifu Chen, Shengpeng Ji, Zhengqing Liu +6
Achieving seamless, human-like interaction remains a key challenge for full-duplex spoken dialogue models (SDMs). Reinforcement learning (RL) has substantially enhanced text- and v…
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
Hao-Xiang Xu, Chong Deng, Jiaqing Liu +5
Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and broad coverage of scenario. Ho…