3 papers
eess.AS2026
ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure
Zixiang Wan, Xusheng Yang, Zheng Wang +1
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to pr…
cs.SD2025
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
Xusheng Yang, Long Zhou, Wenfu Wang +6
We propose \textbf{U-Codec}, an \textbf{U}ltra low frame-rate neural speech \textbf{Codec} that achieves high-fidelity reconstruction and fast speech generation at an extremely low…
eess.AS2025
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
Yushen Chen, Kai Hu, Long Zhou +4
We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, musi…