From the 1 of 6 linked papers with an AI index.
6 papers
Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach
Junkun Chen, Jian Xue, Ming Tang +4
The paper proposes a data‑driven prefix‑to‑prefix fine‑tuning method for simultaneous speech translation that works with decoder‑only LLMs without changing their architecture, usin…
Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving
Ruchao Fan, Yiming Wang, Rui Zhao +10
Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We…
PHRASED: Phrase Dictionary Biasing for Speech Translation
Peidong Wang, Jian Xue, Rui Zhao +3
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in spe…
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Microsoft, :, Abdelrahman Abouelenin +73
We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Peidong Wang, Naoyuki Kanda, Jian Xue +7
Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change o…
Proto-OOD: Enhancing OOD Object Detection with Prototype Feature Similarity
Junkun Chen, Jilin Mei, Liang Chen +3
Neural networks that are trained on limited category samples often mispredict out-of-distribution (OOD) objects. We observe that features of the same category are more tightly clus…