1 citations · 1 across the 4 of their papers we have counts for
5 papers
A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction
Qing Wang, Zehan Li, Yaodong Song +6
This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT)…
Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
Zehan Li, Yan Yang, Xueqing Li +3
Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SS…
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
Hongjie Chen, Zehan Li, Yaodong Song +13
Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing m…
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
Yaodong Song, Hongjie Chen, Jie Lian +8
While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions be…
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Xueqing Li, Hao Ma, Zehan Li +8
Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…