3 papers
cs.AI2025
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
Dongchao Yang, Songxiang Liu, Disong Wang +3
Recent advances in Omni models have enabled unified multimodal perception and generation. However, most existing systems still exhibit rigid reasoning behaviors, either overthinkin…
cs.CL2025
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
Dingdong Wang, Junan Li, Mingyu Cui +3
With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has d…
cs.SD2025
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
Dingdong Wang, Jin Xu, Ruihang Chu +6
Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to s…