From the 1 of 7 linked papers with an AI index.
7 papers
Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach
Junkun Chen, Jian Xue, Ming Tang +4
The paper proposes a data‑driven prefix‑to‑prefix fine‑tuning method for simultaneous speech translation that works with decoder‑only LLMs without changing their architecture, usin…
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation
Yuxuan Hu, Heng Lu, Ruchao Fan +8
Strong speech-to-text (S2T) LLMs already provide robust speech perception and text reasoning, but adding speech-to-speech (S2S) output is challenging: fine-tuning the backbone can…
PHRASED: Phrase Dictionary Biasing for Speech Translation
Peidong Wang, Jian Xue, Rui Zhao +3
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in spe…
Length Aware Speech Translation for Video Dubbing
Harveen Singh Chadha, Aswin Shanmugam Subramanian, Vikas Joshi +4
In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video du…
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Peidong Wang, Naoyuki Kanda, Jian Xue +7
Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change o…
Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages
Midia Yousefi, Yao Qian, Junkun Chen +5
End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applicat…