Sequence-to-Sequence Models Can Directly Translate Foreign Speech
arXiv:1703.08581
Abstract
We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another. The model does not explicitly transcribe the speech into text in the source language, nor does it require supervision from the ground truth source language transcription during training. We apply a slightly modified sequence-to-sequence with attention architecture that has previously been used for speech recognition and show that it can be repurposed for this more complex task, illustrating the power of attention-based models. A single model trained end-to-end obtains state-of-the-art performance on the Fisher Callhome Spanish-English speech translation task, outperforming a cascade of independently trained sequence-to-sequence speech recognition and machine translation models by 1.8 BLEU points on the Fisher test set. In addition, we find that making use of the training data in both languages by multi-task training sequence-to-sequence speech translation and recognition models with a shared encoder network can improve performance by a further 1.4 BLEU points.
5 pages, 1 figure. Interspeech 2017
References in corpus (4)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Towards better decoding and language model integration in sequence to sequence models
- An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages
Cited by in corpus (24)
- How2: A Large-scale Dataset for Multimodal Language Understanding
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
- Data Efficient Direct Speech-to-Text Translation with Modality Agnostic Meta-Learning
- Improving Clinical Predictions through Unsupervised Time Series Representation Learning
- The USTC-NEL Speech Translation system at IWSLT 2018
- Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation
- Unsupervised Word Segmentation from Speech with Attention
- Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
- A small Griko-Italian speech translation corpus
- : Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Convolutional Speech Recognition with Pitch and Voice Quality Features
- Self-Supervised Representations Improve End-to-End Speech Translation
- Investigating the Reordering Capability in CTC-based Non-Autoregressive End-to-End Speech Translation
- LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition
- UWSpeech: Speech to Speech Translation for Unwritten Languages
- Indicatements that character language models learn English morpho-syntactic units and regularities
- Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation
- BSTC: A Large-Scale Chinese-English Speech Translation Dataset
- Fluent Translations from Disfluent Speech in End-to-End Speech Translation
- Cross-modal Spectrum Transformation Network For Acoustic Scene classification
- A Technical Report: BUT Speech Translation Systems
- Towards Fluent Translations from Disfluent Speech