activity
20242026
collaborators

14 papers

cs.CL2026

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

Pengchao Feng, Chao-Hong Tan, Qian Chen +3

Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoke…

eess.AS2026

BareWave: Waveform-Native Flow-Matching Text-to-Speech

Wei Fan, Chao-Hong Tan, Qian Chen +5

Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality syst…

cs.SD2026

UniVocal: Unified Speech-Singing Code-Switching Synthesis

Yufei Shi, Qian Chen, Wen Wang +3

We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…

cs.AI2026

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

Yifu Chen, Shengpeng Ji, Qian Chen +9

End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems.…

cs.AI2026

Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models

Yifu Chen, Shengpeng Ji, Zhengqing Liu +6

Achieving seamless, human-like interaction remains a key challenge for full-duplex spoken dialogue models (SDMs). Reinforcement learning (RL) has substantially enhanced text- and v…

cs.CL2026

GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling

Hao-Xiang Xu, Chong Deng, Jiaqing Liu +5

Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and broad coverage of scenario. Ho…