SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
arXiv:1910.11559
Abstract
While various end-to-end models for spoken language understanding tasks have been explored recently, this paper is probably the first known attempt to challenge the very difficult task of end-to-end spoken question answering (SQA). Learning from the very successful BERT model for various text processing tasks, here we proposed an audio-and-text jointly learned SpeechBERT model. This model outperformed the conventional approach of cascading ASR with the following text question answering (TQA) model on datasets including ASR errors in answer spans, because the end-to-end model was shown to be able to extract information out of audio data before ASR produced errors. When ensembling the proposed end-to-end model with the cascade architecture, even better performance was achieved. In addition to the potential of end-to-end SQA, the SpeechBERT can also be considered for many other spoken language understanding tasks just as BERT for many text processing tasks.
Interspeech 2020
References in corpus (9)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Cross-lingual Language Model Pretraining
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Word Translation Without Parallel Data
- Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
- Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Unsupervised pre-training for sequence to sequence speech recognition
Cited by in corpus (15)
- A Survey on Visual Transformer
- Pre-trained Models for Natural Language Processing: A Survey
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Towards Data Distillation for End-to-end Spoken Conversational Question Answering
- Pre-Trained Models: Past, Present and Future
- MAM: Masked Acoustic Modeling for End-to-End Speech-to-Text Translation
- Speech-language Pre-training for End-to-end Spoken Language Understanding
- Towards Semi-Supervised Semantics Understanding from Speech
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation
- Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining
- Weakly Supervised Construction of ASR Systems with Massive Video Data
- ST-BERT: Cross-modal Language Model Pre-training For End-to-end Spoken Language Understanding
- An Audio-enriched BERT-based Framework for Spoken Multiple-choice Question Answering
- Predicting times of waiting on red signals using BERT
- FANS: Fusing ASR and NLU for on-device SLU