1 paper
Ayush Barik, Sofia Stoica, Nikhil Sarda +4
Text-to-audio diffusion models produce high-fidelity audio but require tens of function evaluations (NFEs), incurring multi-second latency and limited throughput. We present SoundW…