activity
20182022
most citedUnsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages

3 citations · 6 across the 2 of their papers we have counts for

collaborators

8 papers

cs.SD2022

DGC-vector: A new speaker embedding for zero-shot voice conversion

Ruitong Xiao, Haitong Zhang, Yue Lin

Recently, more and more zero-shot voice conversion algorithms have been proposed. As a fundamental part of zero-shot voice conversion, speaker embeddings are the key to improving t…

cs.SD2022

Improve few-shot voice cloning using multi-modal learning

Haitong Zhang, Yue Lin

Recently, few-shot voice cloning has achieved a significant improvement. However, most models for few-shot voice cloning are single-modal, and multi-modal few-shot voice cloning ha…

cs.CL20213 cited

Revisiting IPA-based Cross-lingual Text-to-speech

Haitong Zhang, Haoyue Zhan, Yang Zhang +2

International Phonetic Alphabet (IPA) has been widely used in cross-lingual text-to-speech (TTS) to achieve cross-lingual voice cloning (CL VC). However, IPA itself has been unders…

cs.SD2020

The NeteaseGames System for Voice Conversion Challenge 2020 with Vector-quantization Variational Autoencoder and WaveNet

Haitong Zhang

This paper presents the description of our submitted system for Voice Conversion Challenge (VCC) 2020 with vector-quantization variational autoencoder (VQ-VAE) with WaveNet as the…

eess.AS20203 cited

Unsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages

Haitong Zhang, Yue Lin

Recently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS). These models can generate near-human speech with a large accurately-tra…

cs.CL2019

Improving Interpretability of Word Embeddings by Generating Definition and Usage

Haitong Zhang, Yongping Du, Jiaxin Sun +1

Word embeddings are substantially successful in capturing semantic relations among words. However, these lexical semantics are difficult to be interpreted. Definition modeling prov…