activity
20222024
most citedM2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis

3 citations · 7 across the 9 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2024★ 1 cited

Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining

Jinlong Xue, Yayue Deng, Yingming Gao +1

Recent prompt-based text-to-speech (TTS) models can clone an unseen speaker using only a short speech prompt. They leverage a strong in-context ability to mimic the speech prompts,…

cs.SD2024

Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model

Jinlong Xue, Yayue Deng, Yicheng Han +2

Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speec…

cs.SD2024

Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Jinlong Xue, Yayue Deng, Yingming Gao +1

Recent advancements in diffusion models and large language models (LLMs) have significantly propelled the field of AIGC. Text-to-Audio (TTA), a burgeoning AIGC application designed…

cs.SD2023★ 1 cited

Frame-level emotional state alignment method for speech emotion recognition

Qifei Li, Yingming Gao, Cong Wang +4

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level la…

cs.SD2023★ 3 cited

M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis

Jinlong Xue, Yayue Deng, Fengping Wang +5

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively…

cs.SD2022

ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis

Jinlong Xue, Yayue Deng, Yichen Han +3

In recent years, neural network based methods for multi-speaker text-to-speech synthesis (TTS) have made significant progress. However, the current speaker encoder models used in t…