8 papers
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
Zhengyan Sheng, Zhihao Du, Shiliang Zhang +2
Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models exp…
Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR
Yue Gu, Zhihao Du, Ying Shi +2
Fine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic persona…
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
Yue Gu, Zhihao Du, Ying Shi +3
Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effect…
Differentiable Reward Optimization for LLM based TTS system
Changfeng Gao, Zhihao Du, Shiliang Zhang
This paper proposes a novel Differentiable Reward Optimization (DiffRO) method aimed at enhancing the performance of neural codec language models based text-to-speech (TTS) systems…
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
Guanrou Yang, Fan Yu, Ziyang Ma +4
While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompa…
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Keyu An, Qian Chen, Chong Deng +30
This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative mo…