3 papers
eess.AS2024
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13
Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…
cs.SD2024
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
Haohan Guo, Fenglong Xie, Dongchao Yang +2
The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by ``recency bias", CLM lacks sufficient attentio…
cs.SD2024
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
Haohan Guo, Fenglong Xie, Kun Xie +4
The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered…