activity
20192024
most citedHiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

20 citations · 99 across the 14 of their papers we have counts for

collaborators

20 papers

cs.SD2024

SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Dongchao Yang, Rongjie Huang, Yuanyuan Wang +5

Scaling Text-to-speech (TTS) to large-scale datasets has been demonstrated as an effective method for improving the diversity and naturalness of synthesized speech. At the high lev…

cs.SD2023

UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Dongchao Yang, Jinchuan Tian, Xu Tan +9

Large Language models (LLM) have demonstrated the capability to handle a variety of generative tasks. This paper presents the UniAudio system, which, unlike prior task-specific app…

eess.AS2023

SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias

Sipan Li, Songxiang Liu, Luwen Zhang +5

Generative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small compu…

cs.SD2023

The Singing Voice Conversion Challenge 2023

Wen-Chin Huang, Lester Phillip Violeta, Songxiang Liu +2

We present the latest iteration of the voice conversion challenge (VCC) series, a bi-annual scientific event aiming to compare and understand different voice conversion (VC) system…

cs.SD202320 cited

HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Dongchao Yang, Songxiang Liu, Rongjie Huang +3

Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly…

cs.SD2023

Diverse and Expressive Speech Prosody Prediction with Denoising Diffusion Probabilistic Model

Xiang Li, Songxiang Liu, Max W. Y. Lam +3

Expressive human speech generally abounds with rich and flexible speech prosody variations. The speech prosody predictors in existing expressive speech synthesis methods mostly pro…