activity
20172026
most citedVQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature

43 citations · 170 across the 161 of their papers we have counts for

collaborators
Showing 2024Show all

41 papers · 1 filter

eess.AS2024

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Wenxi Chen, Ziyang Ma, Ruiqi Yan +13

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a…

cs.CV2024

VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization

Tao Liu, Ziyang Ma, Qi Chen +4

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across d…

eess.AS2024★ 1 cited

Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection

Rongfeng Su, Changqing Xu, Xinyi Wu +4

Previous studies have demonstrated that emotional features from a single acoustic sentiment label can enhance depression diagnosis accuracy. Additionally, according to the Emotion…

cs.AI2024

A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

Zheshu Song, Ziyang Ma, Yifan Yang +2

Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in t…

eess.AS2024

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

Hankun Wang, Haoran Wang, Yiwei Guo +3

Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coh…

eess.AS2024

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Yifan Yang, Shujie Liu, Jinyu Li +10

This paper introduces Interleaved Speech-Text Language Model (IST-LM) for zero-shot streaming Text-to-Speech (TTS). Unlike many previous approaches, IST-LM is directly trained on i…