62 citations · 190 across the 31 of their papers we have counts for
4 papers · 1 filter
Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
Yifan Yang, Feiyu Shen, Chenpeng Du +4
Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer…
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
Yiwei Guo, Chenpeng Du, Ziyang Ma +2
Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms th…
Improving Audio Caption Fluency with Automatic Error Correction
Hanxue Zhang, Zeyu Xie, Xuenan Xu +2
Automated audio captioning (AAC) is an important cross-modality translation task, aiming at generating descriptions for audio clips. However, captions generated by previous AAC mod…
DiffVoice: Text-to-Speech with Latent Diffusion
Zhijun Liu, Yiwei Guo, Kai Yu
In this work, we present DiffVoice, a novel text-to-speech model based on latent diffusion. We propose to first encode speech signals into a phoneme-rate latent representation with…