43 citations · 77 across the 7 of their papers we have counts for
7 papers
Improve Bilingual TTS Using Dynamic Language and Phonology Embedding
Fengyu Yang, Jian Luan, Yujun Wang
In most cases, bilingual TTS needs to handle three types of input scripts: first language only, second language only, and second language embedded in the first language. In the lat…
HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
Jiawei Chen, Xu Tan, Jian Luan +2
High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and…
Transfer Learning for Improving Singing-voice Detection in Polyphonic Instrumental Music
Yuanbo Hou, Frank K. Soong, Jian Luan +1
Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-voca…
PPSpeech: Phrase based Parallel End-to-End TTS System
Yahuan Cong, Ran Zhang, Jian Luan
Current end-to-end autoregressive TTS systems (e.g. Tacotron 2) have outperformed traditional parallel approaches on the quality of synthesized speech. However, they introduce new…
DeepSinger: Singing Voice Synthesis with Data Mined From the Web
Yi Ren, Xu Tan, Tao Qin +3
In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music we…
Adversarially Trained Multi-Singer Sequence-To-Sequence Singing Synthesizer
Jie Wu, Jian Luan
This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings. Based on the sequence-to-sequence singing model, we design a…