most citedM2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis

3 citations · 6 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD20241 cited

Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining

Jinlong Xue, Yayue Deng, Yingming Gao +1

Recent prompt-based text-to-speech (TTS) models can clone an unseen speaker using only a short speech prompt. They leverage a strong in-context ability to mimic the speech prompts,…

cs.SD2024

Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model

Jinlong Xue, Yayue Deng, Yicheng Han +2

Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speec…

cs.SD2024

Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Jinlong Xue, Yayue Deng, Yingming Gao +1

Recent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of AIGC. Text-to-Audio (TTA), a burgeoning AIGC application designed…

cs.SD20231 cited

Frame-level emotional state alignment method for speech emotion recognition

Qifei Li, Yingming Gao, Cong Wang +4

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level la…

cs.AI20231 cited

Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis

Dengfeng Ke, Yayue Deng, Yukang Jia +6

Regressive Text-to-Speech (TTS) system utilizes attention mechanism to generate alignment between text and acoustic feature sequence. Alignment determines synthesis robustness (e.g…

cs.SD20233 cited

M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis

Jinlong Xue, Yayue Deng, Fengping Wang +5

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively…