collaborators

5 papers

cs.SD2026

Eliminating stability hallucinations in llm-based tts models via attention guidance

ShiMing Wang, ZhiHao Du, Yang Xiang +6

This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mec…

cs.SD2025

The CCF AATC 2025 Speech Restoration Challenge: A Retrospective

Junan Zhang, Mengyao Zhu, Xin Xu +3

Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, i…

eess.AS2025

Group Relative Policy Optimization for Text-to-Speech with Large Language Models

Chang Liu, Ya-Jun Hu, Ying-Ying Gao +2

This paper proposes a GRPO-based approach to enhance the performance of large language model (LLM)-based text-to-speech (TTS) models by deriving rewards from an off-the-shelf autom…

eess.AS2025

Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models

Jing-Xuan Zhang, Genshun Wan, Jianqing Gao +1

Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundatio…

cs.SD2025

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

Zhengyan Sheng, Zhihao Du, Heng Lu +2

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generat…