6 citations · 7 across the 12 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
Yakun Song, Xiaobin Zhuang, Jiawei Chen +8
Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under d…
eess.AS2025
AudioMorphix: Training-free audio editing with diffusion probabilistic models
Jinhua Liang, Yuanzhe Chen, Yi Yuan +5
Editing sound with precision is a crucial yet underexplored challenge in audio content creation. While existing works can manipulate sounds by text instructions or audio exemplar p…
eess.AS2025
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
Dongya Jia, Zhuo Chen, Jiawei Chen +8
Several recent studies have attempted to autoregressively generate continuous speech representations without discrete speech tokens by combining diffusion and autoregressive models…