most citedHiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

43 citations · 77 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD20221 cited

Improve Bilingual TTS Using Dynamic Language and Phonology Embedding

Fengyu Yang, Jian Luan, Yujun Wang

In most cases, bilingual TTS needs to handle three types of input scripts: first language only, second language only, and second language embedded in the first language. In the lat…

eess.AS202043 cited

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

Jiawei Chen, Xu Tan, Jian Luan +2

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and…

eess.AS2020

Transfer Learning for Improving Singing-voice Detection in Polyphonic Instrumental Music

Yuanbo Hou, Frank K. Soong, Jian Luan +1

Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-voca…

eess.AS20201 cited

PPSpeech: Phrase based Parallel End-to-End TTS System

Yahuan Cong, Ran Zhang, Jian Luan

Current end-to-end autoregressive TTS systems (e.g. Tacotron 2) have outperformed traditional parallel approaches on the quality of synthesized speech. However, they introduce new…

eess.AS20204 cited

DeepSinger: Singing Voice Synthesis with Data Mined From the Web

Yi Ren, Xu Tan, Tao Qin +3

In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music we…

eess.AS20208 cited

Adversarially Trained Multi-Singer Sequence-To-Sequence Singing Synthesizer

Jie Wu, Jian Luan

This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings. Based on the sequence-to-sequence singing model, we design a…