37 citations · 68 across the 11 of their papers we have counts for
5 papers · 1 filter
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis
Yu Gu, Yianrao Bian, Guangzhi Lei +2
This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original…
SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias
Sipan Li, Songxiang Liu, Luwen Zhang +5
Generative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small compu…
Complexity Scaling for Speech Denoising
Hangting Chen, Jianwei Yu, Chao Weng
Computational complexity is critical when deploying deep learning-based speech denoising models for on-device applications. Most prior research focused on optimizing model architec…
Make-A-Voice: Unified Voice Synthesis With Discrete Representation
Rongjie Huang, Chunlei Zhang, Yongqi Wang +7
Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…
Cross-Age Speaker Verification: Learning Age-Invariant Speaker Embeddings
Xiaoyi Qin, Na Li, Chao Weng +2
Automatic speaker verification has achieved remarkable progress in recent years. However, there is little research on cross-age speaker verification (CASV) due to insufficient rele…