Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
Tao Li, Wenshuo Ge, Zhichao Wang +6
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-…
cs.SD2025
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
Xiaoxue Luo, Jinwei Huang, Runyan Yang +4
Universal audio codecs learn entangled representations across audio types, whereas some specific codecs offer decoupled representations but are limited to speech. Real-world audio,…