11 papers
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
Shuhao Xu, Yifan Hu, Jingjing Wu +3
Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grai…
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
Cong Wang, Changfeng Gao, Yang Xiang +7
Differentiable reinforcement learning (RL) frameworks like DiffRO offer a powerful approach for controllable text-to-speech (TTS), but are vulnerable to reward hacking, particularl…
Eliminating stability hallucinations in llm-based tts models via attention guidance
ShiMing Wang, ZhiHao Du, Yang Xiang +6
This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mec…
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
Keyu An, Zhiyu Zhang, Changfeng Gao +7
This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram fram…
Explore the Reinforcement Learning for the LLM based ASR and TTS system
Changfeng Gao, Yabin Li, Keyu An +4
In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL…
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
Guanrou Yang, Chen Yang, Qian Chen +12
Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made h…