7 papers
Depth-Entropy Guided Sampling for Training-Free LLM Reasoning
Zibin Meng, Peng Xie, Kani Chen
Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, an…
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
Guojian Li, Zhixian Zhao, Zhennan Lin +9
While speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is eviden…
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
Huakang Chen, Jingbin Hu, Liumeng Xue +12
Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due t…
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
Zhennan Lin, Shuai Wang, Zhaokai Sun +5
Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker t…
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
Jingbin Hu, Haoyu Zhang, Dake Guo +10
Large Language Models (LLMs) have advanced audio generation through discrete representation learning. However, most existing neural codecs focus on speech and emphasize reconstruct…
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
Peng Xie, Xingyuan Liu, Tsz Wai Chan +5
Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phe…