4 papers
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
Raphaël Bagat, Zhe Zhang, Junichi Yamagishi +2
Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (…
BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
Raphaël Bagat, Irina Illina, Emmanuel Vincent
Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder…
Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
Raphaël Bagat, Irina Illina, Emmanuel Vincent
We aim to improve the robustness of Automatic Speech Recognition (ASR) systems against non-native speech, particularly in low-resourced multi-accent settings. We introduce Mixture…
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
Sewade Ogun, Vincent Colotte, Emmanuel Vincent
Augmenting the training data of automatic speech recognition (ASR) systems with synthetic data generated by text-to-speech (TTS) or voice conversion (VC) has gained popularity in r…