natural language processing

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

arXiv:2607.13158

summary

The paper proposes a data‑driven prefix‑to‑prefix fine‑tuning method for simultaneous speech translation that works with decoder‑only LLMs without changing their architecture, using fixed‑length chunks and a rewind‑based committed prefix to improve translation quality at low latency.

Abstract

Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, yet remains challenging for decoder-only LLM systems due to limited context and cross-lingual reordering. Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech where segmentation boundaries are ambiguous. We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled prefix-to-prefix (P2P) targets with bounded waiting for fine-tuning, yielding CSSEL-P2P, where CSSEL is our proposed chunked streaming speech encoder LLM. In our in-house conversational speech evaluation, CSSEL-P2P improves streaming quality by +1.54 COMETKiwi over the CSSEL streaming baseline at comparable latency (+0.15s Average Lagging), suggesting effective SimulST without architectural changes via P2P supervision.

Topics & keywords

#simultaneous speech translation#prefix-to-prefix supervision#streaming decoding#language models#speech encoderprefix-to-prefixchunked streaming speech encoderrewind-based committed prefixCOMETKiwiaverage laggingteacher-labeled targets