activity
20182022
most citedVoice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

6 citations · 6 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2022

Neural Grapheme-to-Phoneme Conversion with Pre-trained Grapheme Models

Lu Dong, Zhi-Qiang Guo, Chao-Hong Tan +3

Neural network models have achieved state-of-the-art performance on grapheme-to-phoneme (G2P) conversion. However, their performance relies on large-scale pronunciation dictionarie…

eess.AS20206 cited

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen +4

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR a…

eess.AS2019

ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech

Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37

Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…

eess.AS2019

End-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training

Peng-fei Wu, Zhen-hua Ling, Li-juan Liu +3

This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST…

cs.SD2018

Improving Sequence-to-Sequence Acoustic Modeling by Adding Text-Supervision

Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang +3

This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-f…

cs.SD2018

Sequence-to-Sequence Acoustic Modeling for Voice Conversion

Jing-Xuan Zhang, Zhen-Hua Ling, Li-Juan Liu +2

In this paper, a neural network named Sequence-to-sequence ConvErsion NeTwork (SCENT) is presented for acoustic modeling in voice conversion. At training stage, a SCENT model is es…