activity
20192024
most citedA Modularized Neural Network with Language-Specific Output Layers for Cross-lingual Voice Conversion

2 citations · 4 across the 3 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20241 cited

RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging

Mingyang Zhang, Yi Zhou, Yi Ren +3

This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally t…

eess.AS2024

Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis

Xuehao Zhou, Mingyang Zhang, Yi Zhou +2

Generating speech across different accents while preserving speaker identity is crucial for various real-world applications. However, accurately and independently modeling both spe…

eess.AS20231 cited

Accented Text-to-Speech Synthesis with Limited Data

Xuehao Zhou, Mingyang Zhang, Yi Zhou +2

This paper presents an accented text-to-speech (TTS) synthesis framework with limited training data. We study two aspects concerning accent rendering: phonetic (phoneme difference)…

eess.AS2020

Transfer Learning from Speech Synthesis to Voice Conversion with Non-Parallel Training Data

Mingyang Zhang, Yi Zhou, Li Zhao +1

This paper presents a novel framework to build a voice conversion (VC) system by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning. We…

eess.AS20192 cited

A Modularized Neural Network with Language-Specific Output Layers for Cross-lingual Voice Conversion

Yi Zhou, Xiaohai Tian, Emre Yılmaz +2

This paper presents a cross-lingual voice conversion framework that adopts a modularized neural network. The modularized neural network has a common input structure that is shared…