4 papers
Using multiple reference audios and style embedding constraints for speech synthesis
Cheng Gong, Longbiao Wang, Zhenhua Ling +2
The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the r…
Information Sieve: Content Leakage Reduction in End-to-End Prosody For Expressive Speech Synthesis
Xudong Dai, Cheng Gong, Longbiao Wang +1
Expressive neural text-to-speech (TTS) systems incorporate a style encoder to learn a latent embedding as the style information. However, this embedding process may encode redundan…
DiDiSpeech: A Large Scale Mandarin Speech Corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang +8
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…
DELTA: A DEep learning based Language Technology plAtform
Kun Han, Junwen Chen, Hui Zhang +20
In this paper we present DELTA, a deep learning based language technology platform. DELTA is an end-to-end platform designed to solve industry level natural language and speech pro…