most citedGOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2026

A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction

Qing Wang, Zehan Li, Yaodong Song +6

This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT)…

eess.AS2025

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy

Zehan Li, Yan Yang, Xueqing Li +3

Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SS…

cs.CL2025

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Hongjie Chen, Zehan Li, Yaodong Song +13

Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing m…

cs.CL20251 cited

GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM

Yaodong Song, Hongjie Chen, Jie Lian +8

While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions be…

eess.AS2025

Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization

Xueqing Li, Hao Ma, Zehan Li +8

Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…