From the 1 of 4 linked papers with an AI index.
4 papers
Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach
Junkun Chen, Jian Xue, Ming Tang +4
The paper proposes a data‑driven prefix‑to‑prefix fine‑tuning method for simultaneous speech translation that works with decoder‑only LLMs without changing their architecture, usin…
PHRASED: Phrase Dictionary Biasing for Speech Translation
Peidong Wang, Jian Xue, Rui Zhao +3
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in spe…
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Peidong Wang, Naoyuki Kanda, Jian Xue +7
Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change o…
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit +4
Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that G…