activity
20192024
most citedSemi-supervised Learning for Multi-speaker Text-to-speech Synthesis Using Discrete Speech Representation

2 citations · 4 across the 5 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2024

StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion

Zhichao Wang, Yuanzhe Chen, Xinsheng Wang +2

StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to…

eess.AS2024

StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion

Zhichao Wang, Yuanzhe Chen, Xinsheng Wang +2

Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion…

eess.AS2023

LM-VC: Zero-shot Voice Conversion via Speech Generation based on Language Models

Zhichao Wang, Yuanzhe Chen, Lei Xie +2

Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation. In this paper, we explo…

eess.AS2023

Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion

Zhichao Wang, Liumeng Xue, Qiuqiang Kong +4

Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typ…

eess.AS2022

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

Yuanzhe Chen, Ming Tu, Tang Li +7

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracte…

eess.AS2021

Cloning one's voice using very limited data in the wild

Dongyang Dai, Yuanzhe Chen, Li Chen +6

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily a…