6 papers
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
Zhiwei Lin, Jun Chen, Boshi Tang +5
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous…
STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity
Sitong Cheng, Weizhen Bian, Songjun Cao +9
Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue),…
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
Weihao Wu, Liang Cao, Xinyu Wu +4
Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create im…
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
Rui Niu, Weihao Wu, Jie Chen +2
Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with rob…
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Weihao wu, Zhiwei Lin, Yixuan Zhou +6
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
Shuoyi Zhou, Yixuan Zhou, Weiqin Li +6
This paper describes the zero-shot spontaneous style TTS system for the ISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose a LLaMA-based codec language model with…