8 citations · 8 across the 6 of their papers we have counts for
1 paper · 1 filter
Yakun Song, Zhuo Chen, Xiaofei Wang +3
Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference s…