most citedNaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

38 citations · 52 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2023

Illumination Controllable Dehazing Network based on Unsupervised Retinex Embedding

Jie Gui, Xiaofeng Cong, Lei He +2

On the one hand, the dehazing task is an illposedness problem, which means that no unique solution exists. On the other hand, the dehazing task should take into account the subject…

eess.AS202338 cited

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Kai Shen, Zeqian Ju, Xu Tan +6

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, an…

cs.NE20234 cited

A Reinforcement Learning-assisted Genetic Programming Algorithm for Team Formation Problem Considering Person-Job Matching

Yangyang Guo, Hao Wang, Lei He +3

An efficient team is essential for the company to successfully complete new projects. To solve the team formation problem considering person-job matching (TFP-PJM), a 0-1 integer p…

cs.SD20239 cited

AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models

Yuancheng Wang, Zeqian Ju, Xu Tan +4

Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-bas…

cs.SD20221 cited

DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders

Yanqing Liu, Ruiqing Xue, Lei He +2

Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from tw…

cs.CV2022

ReMix: A General and Efficient Framework for Multiple Instance Learning based Whole Slide Image Classification

Jiawei Yang, Hanbo Chen, Yu Zhao +4

Whole slide image (WSI) classification often relies on deep weakly supervised multiple instance learning (MIL) methods to handle gigapixel resolution images and slide-level labels.…