4 papers
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Wenrui Liu, Qian Chen, Wen Wang +11
Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neu…
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
Dan Luo, Chengyuan Ma, Weiqin Li +3
With the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the…
UniSep: Universal Target Audio Separation with Language Models at Scale
Yuanyuan Wang, Hangting Chen, Dongchao Yang +7
We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep…
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
Shuoyi Zhou, Yixuan Zhou, Weiqin Li +6
This paper describes the zero-shot spontaneous style TTS system for the ISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose a LLaMA-based codec language model with…