activity
20242026
most citedWavChat: A Survey of Spoken Dialogue Models

2 citations · 2 across the 7 of their papers we have counts for

collaborators

11 papers

cs.SD2026

A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation

Hanchen Pei, Shujie Liu, Yanqing Liu +5

Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody,…

cs.CL2025

TrInk: Ink Generation with Transformer Network

Zezhong Jin, Shubhang Desai, Xu Chen +8

In this paper, we propose TrInk, a Transformer-based model for ink generation, which effectively captures global dependencies. To better facilitate the alignment between the input…

cs.SD2025

Next Tokens Denoising for Speech Synthesis

Yanqing Liu, Ruiqing Xue, Chong Zhang +7

While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, c…

cs.SD2025

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

Hui Wang, Yifan Yang, Shujie Liu +7

Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time appli…

cs.LG2025

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Haiyang Sun, Shujie Hu, Shujie Liu +8

Zero-shot streaming text-to-speech is an important research topic in human-computer interaction. Existing methods primarily use a lookahead mechanism, relying on future text to ach…

cs.LG2025

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

Peng Xia, Jinglu Wang, Yibo Peng +10

Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across div…