3 papers
cs.CL2026
Speech LLMs are Contextual Reasoning Transcribers
Keqi Deng, Ruchao Fan, Bo Ren +2
Despite extensions to speech inputs, effectively leveraging the rich knowledge and contextual understanding of large language models (LLMs) in automatic speech recognition (ASR) re…
eess.AS2025
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
Mohan Shi, Xiong Xiao, Ruchao Fan +2
Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end spe…
eess.AS2025
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
Shaoshi Ling, Gang Liu, Guoli Ye +1
Speech summarization is a critical component of spoken content understanding, particularly in the era of rapidly growing spoken and audiovisual data. Recent advances in multi-modal…