activity
20242026
collaborators

7 papers

cs.SD2026

AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models

Wenyu Li, Xiaoqi Jiao, Yi Chang +2

The creation of high-quality multimodal datasets remains fundamental for advancing role-playing capabilities in large language models (LLMs). While existing works predominantly foc…

cs.CL2025

Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigm

Zhuo Li, Yuhao Du, Xiaoqi Jiao +5

Selecting high-quality and diverse training samples from extensive datasets plays a crucial role in reducing training overhead and enhancing the performance of Large Language Model…

cs.CL2025

Entropy-based Coarse and Compressed Semantic Speech Representation Learning

Jialong Zuo, Guangyan Zhang, Minghui Fang +5

Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms int…

cs.CL2025

Recent Advances in Speech Language Models: A Survey

Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao +5

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often rel…

eess.AS2025

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

Zhuangfei Cheng, Guangyan Zhang, Zehai Tu +6

Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, t…

cs.CL2025

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Wenqian Cui, Xiaoqi Jiao, Ziqiao Meng +1

With the rising need for speech-based interaction models, end-to-end Spoken Language Models (SLMs) have emerged as a promising solution. While these models require comprehensive wo…