38 citations · 78 across the 12 of their papers we have counts for
12 papers
CoMoSVC: Consistency Model-based Singing Voice Conversion
Yiwen Lu, Zhen Ye, Wei Xue +3
The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre. However, the i…
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Kai Shen, Zeqian Ju, Xu Tan +6
Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, an…
AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models
Yuancheng Wang, Zeqian Ju, Xu Tan +4
Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-bas…
Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation
Qi Chen, Ziyang Ma, Tao Liu +4
Audio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive…
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
Ruiqing Xue, Yanqing Liu, Lei He +4
Neural text-to-speech (TTS) generally consists of cascaded architecture with separately optimized acoustic model and vocoder, or end-to-end architecture with continuous mel-spectro…
A Study on ReLU and Softmax in Transformer
Kai Shen, Junliang Guo, Xu Tan +3
The Transformer architecture consists of self-attention and feed-forward networks (FFNs) which can be viewed as key-value memories according to previous works. However, FFN and tra…