20 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.SD2024
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
Dongchao Yang, Rongjie Huang, Yuanyuan Wang +5
Scaling Text-to-speech (TTS) to large-scale datasets has been demonstrated as an effective method for improving the diversity and naturalness of synthesized speech. At the high lev…
eess.AS2023
SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias
Sipan Li, Songxiang Liu, Luwen Zhang +5
Generative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small compu…
cs.SD2023★ 20 cited
HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Dongchao Yang, Songxiang Liu, Rongjie Huang +3
Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly…