8 papers
When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation
Zeyu Yang, Satoshi Nakamura
Extra context is valuable for simultaneous speech translation of technical talks, but injecting the entire document context into every streaming segment is often too coarse. Throug…
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
Yuchen Song, Xi Chen, Mingze Li +1
Speech-to-speech translation (S2ST) systems have achieved impressive progress in semantic accuracy and speech naturalness. However, the cross-lingual transfer of lexical stress, a…
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
Qixu Chen, Satoshi Nakamura
Large-scale mined corpora provide abundant training data for end-to-end speech-to-speech translation (S2ST) but may contain noise, misalignment, and semantic errors. Filtering nois…
Gradient-Informed Training for Low-Resource Multilingual Speech Translation
Ruiyan Sun, Satoshi Nakamura
In low-resource multilingual speech-to-text translation, uniform architectural sharing across languages frequently introduces representation conflicts that impede convergence. This…
Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies
Qianen Zhang, Zeyu Yang, Satoshi Nakamura
Simultaneous Machine Translation (SiMT) requires high-quality translations under strict real-time constraints, which traditional policies with only READ/WRITE actions cannot fully…
DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation
Zeyu Yang, Satoshi Nakamura
Simultaneous speech translation requires accurate segmentation to balance translation quality and latency. Recent studies such as SHAS have introduced pretrained segmentation model…