1 paper · 1 filter
Detai Xin, Shujie Hu, Chengzuo Yang +4
We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that…