43 citations · 170 across the 161 of their papers we have counts for
41 papers · 1 filter
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
Wenxi Chen, Ziyang Ma, Ruiqi Yan +13
Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a…
VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
Tao Liu, Ziyang Ma, Qi Chen +4
We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across d…
Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection
Rongfeng Su, Changqing Xu, Xinyi Wu +4
Previous studies have demonstrated that emotional features from a single acoustic sentiment label can enhance depression diagnosis accuracy. Additionally, according to the Emotion…
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
Zheshu Song, Ziyang Ma, Yifan Yang +2
Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in t…
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
Hankun Wang, Haoran Wang, Yiwei Guo +3
Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coh…
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Yifan Yang, Shujie Liu, Jinyu Li +10
This paper introduces Interleaved Speech-Text Language Model (IST-LM) for zero-shot streaming Text-to-Speech (TTS). Unlike many previous approaches, IST-LM is directly trained on i…