Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation
arXiv:2309.02459 · doi:10.21437/Interspeech.2023-1378
Abstract
Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) performance in new domains. However, the length of speech representation and text representation is inconsistent. Although the previous method up-samples the text representation to align with acoustic modality, it may not match the expected actual duration. In this paper, we proposed novel representations match strategy through down-sampling acoustic representation to align with text modality. By introducing a continuous integrate-and-fire (CIF) module generating acoustic representations consistent with token length, our ASR model can learn unified representations from both modalities better, allowing for domain adaptation using text-only data of the target domain. Experiment results of new domain data demonstrate the effectiveness of the proposed method.
Proceedings of Interspeech. arXiv admin note: text overlap with arXiv:2309.01437
References in corpus (5)
- Sequence Transduction with Recurrent Neural Networks
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
- Improving Hybrid CTC/Attention End-to-end Speech Recognition with Pretrained Acoustic and Language Model
- Improving Mandarin Speech Recogntion with Block-augmented Transformer