activity
20242026
collaborators

8 papers

cs.SD2026

SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer

Zhengyan Sheng, Zhihao Du, Shiliang Zhang +2

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models exp…

cs.SD2025

Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR

Yue Gu, Zhihao Du, Ying Shi +2

Fine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic persona…

cs.CL2025

Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling

Yue Gu, Zhihao Du, Ying Shi +3

Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effect…

cs.SD2025

Differentiable Reward Optimization for LLM based TTS system

Changfeng Gao, Zhihao Du, Shiliang Zhang

This paper proposes a novel Differentiable Reward Optimization (DiffRO) method aimed at enhancing the performance of neural codec language models based text-to-speech (TTS) systems…

eess.AS2024

Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

Guanrou Yang, Fan Yu, Ziyang Ma +4

While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompa…

cs.SD2024

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Keyu An, Qian Chen, Chong Deng +30

This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative mo…