Does Simultaneous Speech Translation need Simultaneous Models?
arXiv:2204.03783 · doi:10.18653/v1/2022.findings-emnlp.11
Abstract
In simultaneous speech translation (SimulST), finding the best trade-off between high translation quality and low latency is a challenging task. To meet the latency constraints posed by the different application scenarios, multiple dedicated SimulST models are usually trained and maintained, generating high computational costs. In this paper, motivated by the increased social and environmental impact caused by these costs, we investigate whether a single model trained offline can serve not only the offline but also the simultaneous task without the need for any additional training or adaptation. Experiments on en->{de, es} indicate that, aside from facilitating the adoption of well-established offline techniques and architectures without affecting latency, the offline solution achieves similar or better translation quality compared to the same model trained in simultaneous settings, as well as being competitive with the SimulST state of the art.
Findings of EMNLP 2022
References in corpus (10)
- Distilling the Knowledge in a Neural Network
- Sequence Transduction with Recurrent Neural Networks
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- fairseq S2T: Fast Speech-to-Text Modeling with fairseq
- Monotonic Multihead Attention
- Bridging the Modality Gap for Speech-to-Text Translation
- SimulMT to SimulST: Adapting Simultaneous Text Translation to End-to-End Simultaneous Speech Translation
- Over-Generation Cannot Be Rewarded: Length-Adaptive Average Lagging for Simultaneous Speech Translation
- Efficient yet Competitive Speech Translation: FBK@IWSLT2022
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring