4 papers
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
Jiacheng Xu, Heting Gao, Liufei Xie +8
Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize…
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
Qi Wang, Zhexu Shen, Meng Chen +4
Vocal-to-accompaniment (V2A) generation, which aims to transform a raw vocal recording into a fully arranged accompaniment, inherently requires jointly addressing an accompaniment…
TQCodec: Towards neural audio codec for high-fidelity music streaming
Lixing He, Zhouxuan Chen, Mingshuai Liu +6
We propose TQCodec, a neural audio codec designed for high-bitrate, high-fidelity music streaming. Unlike existing neural codecs that primarily target ultra-low bitrates (<= 16kbps…
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
Ziqi Dai, Yiting Chen, Jiacheng Xu +8
The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, sc…