12 citations · 20 across the 4 of their papers we have counts for
4 papers · 1 filter
DeepA: A Deep Neural Analyzer For Speech And Singing Vocoding
Sergey Nikonorov, Berrak Sisman, Mingyang Zhang +1
Conventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under…
Transfer Learning from Speech Synthesis to Voice Conversion with Non-Parallel Training Data
Mingyang Zhang, Yi Zhou, Li Zhao +1
This paper presents a novel framework to build a voice conversion (VC) system by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning. We…
Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet
Mingyang Zhang, Xin Wang, Fuming Fang +2
We investigated the training of a shared model for both text-to-speech (TTS) and voice conversion (VC) tasks. We propose using an extended model architecture of Tacotron, that is a…
Error Reduction Network for DBLSTM-based Voice Conversion
Mingyang Zhang, Berrak Sisman, Sai Sirisha Rallabandi +2
So far, many of the deep learning approaches for voice conversion produce good quality speech by using a large amount of training data. This paper presents a Deep Bidirectional Lon…