13 citations · 41 across the 7 of their papers we have counts for
4 papers · 1 filter
ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech
Yi Ren, Ming Lei, Zhiying Huang +4
Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted p…
DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech
Zhiying Huang, Hao Li, Ming Lei
With the number of smart devices increasing, the demand for on-device text-to-speech (TTS) increases rapidly. In recent years, many prominent End-to-End TTS methods have been propo…
Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
Shiliang Zhang, Ming Lei, Zhijie Yan
Connectionist Temporal Classification (CTC) based end-to-end speech recognition system usually need to incorporate an external language model by using WFST-based decoding in order…
Linear networks based speaker adaptation for speech synthesis
Zhiying Huang, Heng Lu, Ming Lei +1
Speaker adaptation methods aim to create fair quality synthesis speech voice font for target speakers while only limited resources available. Recently, as deep neural networks base…