works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CL2026

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

Junkun Chen, Jian Xue, Ming Tang +4

The paper proposes a data‑driven prefix‑to‑prefix fine‑tuning method for simultaneous speech translation that works with decoder‑only LLMs without changing their architecture, usin…

cs.CL2026

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao +10

Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We…

cs.CL2025

PHRASED: Phrase Dictionary Biasing for Speech Translation

Peidong Wang, Jian Xue, Rui Zhao +3

Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in spe…

cs.CL2025

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Microsoft, :, Abdelrahman Abouelenin +73

We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…

cs.SD2025

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation

Peidong Wang, Naoyuki Kanda, Jian Xue +7

Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change o…

cs.CV2025

Proto-OOD: Enhancing OOD Object Detection with Prototype Feature Similarity

Junkun Chen, Jilin Mei, Liang Chen +3

Neural networks that are trained on limited category samples often mispredict out-of-distribution (OOD) objects. We observe that features of the same category are more tightly clus…