activity
20212024
most citedVideoCrafter1: Open Diffusion Models for High-Quality Video Generation

37 citations · 68 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20231 cited

DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis

Yu Gu, Yianrao Bian, Guangzhi Lei +2

This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original…

eess.AS2023

SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias

Sipan Li, Songxiang Liu, Luwen Zhang +5

Generative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small compu…

eess.AS2023

Complexity Scaling for Speech Denoising

Hangting Chen, Jianwei Yu, Chao Weng

Computational complexity is critical when deploying deep learning-based speech denoising models for on-device applications. Most prior research focused on optimizing model architec…

eess.AS20232 cited

Make-A-Voice: Unified Voice Synthesis With Discrete Representation

Rongjie Huang, Chunlei Zhang, Yongqi Wang +7

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…

eess.AS2022

Cross-Age Speaker Verification: Learning Age-Invariant Speaker Embeddings

Xiaoyi Qin, Na Li, Chao Weng +2

Automatic speaker verification has achieved remarkable progress in recent years. However, there is little research on cross-age speaker verification (CASV) due to insufficient rele…