activity
20242026
collaborators

6 papers

cs.CL2026

Speech LLMs are Contextual Reasoning Transcribers

Keqi Deng, Ruchao Fan, Bo Ren +2

Despite extensions to speech inputs, effectively leveraging the rich knowledge and contextual understanding of large language models (LLMs) in automatic speech recognition (ASR) re…

eess.AS2026

RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models

Bo Ren, Ruchao Fan, Yelong Shen +2

Speech large language models (LLMs) have driven significant progress in end-to-end speech understanding and recognition, yet they continue to struggle with accurately recognizing r…

cs.NI2025

Analyzing Communication Predictability in LLM Training

Wenxue Li, Xiangzhou Liu, Yuxuan Li +9

Effective communication is essential in distributed training, with predictability being one of its most significant characteristics. However, existing studies primarily focus on ex…

eess.AS2025

Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model

Haibin Wu, Yuxuan Hu, Ruchao Fan +8

Speech language models (Speech LMs) enable end-to-end speech-text modeling within a single model, offering a promising direction for spoken dialogue systems. The choice of speech-t…

cs.CL2025

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Microsoft, :, Abdelrahman Abouelenin +73

We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…

eess.AS2024

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

Ruchao Fan, Bo Ren, Yuxuan Hu +3

Integrating speech into LLM (speech-LLM) has gaining increased attention recently. The mainstream solution is to connect a well-trained speech encoder and LLM with a neural adapter…