Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
Scaling Speech Tokenizers with Diffusion Autoencoders
Yuancheng Wang, Zhenyu Tang, Yun Wang +9
Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understandi…
cs.SD2025
Overview of the Amphion Toolkit (v0.2)
Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…
cs.SD2025
USED: Universal Speaker Extraction and Diarization
Junyi Ao, Mehmet Sinan Yıldırım, Ruijie Tao +4
Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture…