1 paper
Binh Mai, Tran Quoc Bao Le, Hung Dinh +1
Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step appr…