Magic dust for cross-lingual adaptation of monolingual wav2vec-2.0
arXiv:2110.03560 · doi:10.1109/ICASSP43922.2022.9746276
Abstract
We propose a simple and effective cross-lingual transfer learning method to adapt monolingual wav2vec-2.0 models for Automatic Speech Recognition (ASR) in resource-scarce languages. We show that a monolingual wav2vec-2.0 is a good few-shot ASR learner in several languages. We improve its performance further via several iterations of Dropout Uncertainty-Driven Self-Training (DUST) by using a moderate-sized unlabeled speech dataset in the target language. A key finding of this work is that the adapted monolingual wav2vec-2.0 achieves similar performance as the topline multilingual XLSR model, which is trained on fifty-three languages, on the target language ASR task.
References in corpus (9)
- A Simple Framework for Contrastive Learning of Visual Representations
- Sequence Transduction with Recurrent Neural Networks
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Self-Training for End-to-End Speech Recognition
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Revisiting Self-Training for Neural Sequence Generation
- Iterative Pseudo-Labeling for Speech Recognition
- Exploiting Adapters for Cross-lingual Low-resource Speech Recognition
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans