6 papers
Speech LLMs are Contextual Reasoning Transcribers
Keqi Deng, Ruchao Fan, Bo Ren +2
Despite extensions to speech inputs, effectively leveraging the rich knowledge and contextual understanding of large language models (LLMs) in automatic speech recognition (ASR) re…
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
Bo Ren, Ruchao Fan, Yelong Shen +2
Speech large language models (LLMs) have driven significant progress in end-to-end speech understanding and recognition, yet they continue to struggle with accurately recognizing r…
Analyzing Communication Predictability in LLM Training
Wenxue Li, Xiangzhou Liu, Yuxuan Li +9
Effective communication is essential in distributed training, with predictability being one of its most significant characteristics. However, existing studies primarily focus on ex…
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
Haibin Wu, Yuxuan Hu, Ruchao Fan +8
Speech language models (Speech LMs) enable end-to-end speech-text modeling within a single model, offering a promising direction for spoken dialogue systems. The choice of speech-t…
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Microsoft, :, Abdelrahman Abouelenin +73
We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
Ruchao Fan, Bo Ren, Yuxuan Hu +3
Integrating speech into LLM (speech-LLM) has gaining increased attention recently. The mainstream solution is to connect a well-trained speech encoder and LLM with a neural adapter…