Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation
arXiv:1612.01744
Abstract
This paper proposes a first attempt to build an end-to-end speech-to-text translation system, which does not use source language transcription during learning or decoding. We propose a model for direct speech-to-text translation, which gives promising results on a small French-English synthetic corpus. Relaxing the need for source language transcription would drastically change the data collection methodology in speech translation, especially in under-resourced scenarios. For instance, in the former project DARPA TRANSTAC (speech translation from spoken Arabic dialects), a large effort was devoted to the collection of speech transcripts (and a prerequisite to obtain transcripts was often a detailed transcription guide for languages with little standardized spelling). Now, if end-to-end approaches for speech-to-text translation are successful, one might consider collecting data by asking bilingual speakers to directly utter speech in the source language from target language text utterances. Such an approach has the advantage to be applicable to any unwritten (source) language.
accepted to NIPS workshop on End-to-end Learning for Speech and Audio Processing
Cited by in corpus (34)
- fairseq S2T: Fast Speech-to-Text Modeling with fairseq
- CoVoST 2 and Massively Multilingual Speech-to-Text Translation
- Direct Speech-to-image Translation
- Bridging the Modality Gap for Speech-to-Text Translation
- End-to-End Speech Translation with Knowledge Distillation
- Multilingual Speech Translation with Efficient Finetuning of Pretrained Models
- On Using SpecAugment for End-to-End Speech Translation
- Direct speech-to-speech translation with a sequence-to-sequence model
- Improving Speech Translation by Cross-Modal Multi-Grained Contrastive Learning
- MAM: Masked Acoustic Modeling for End-to-End Speech-to-Text Translation
- Curriculum Pre-training for End-to-End Speech Translation
- Neural Machine Translation: Challenges, Progress and Future
- ESPnet-ST: All-in-One Speech Translation Toolkit
- End-to-end Speech Translation via Cross-modal Progressive Training
- Learning Shared Semantic Space for Speech-to-Text Translation
- Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders
- NeurST: Neural Speech Translation Toolkit
- Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation
- Dealing with training and test segmentation mismatch: FBK@IWSLT2021
- A Data Efficient End-To-End Spoken Language Understanding Architecture
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- UWSpeech: Speech to Speech Translation for Unwritten Languages
- Adaptive Feature Selection for End-to-End Speech Translation
- Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
- Speech-to-speech Translation between Untranscribed Unknown Languages
- On Target Segmentation for Direct Speech Translation
- A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks
- Simultaneous Speech Translation for Live Subtitling: from Delay to Display
- RealTranS: End-to-End Simultaneous Speech Translation with Convolutional Weighted-Shrinking Transformer
- Efficient Transformer for Direct Speech Translation
- Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation
- The NiuTrans End-to-End Speech Translation System for IWSLT 2021 Offline Task
- Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings
- AlloST: Low-resource Speech Translation without Source Transcription