3 papers
cs.SD2025
Next Tokens Denoising for Speech Synthesis
Yanqing Liu, Ruiqing Xue, Chong Zhang +7
While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, c…
cs.SD2025
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
Zhaoxi Mu, Xinyu Yang, Gang Wang
While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, inclu…
cs.CL2024
Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages
Midia Yousefi, Yao Qian, Junkun Chen +5
End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applicat…