Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization
arXiv:2304.14535 · doi:10.1016/j.knosys.2023.110851
Abstract
Automatic speech recognition (ASR) has recently become an important challenge when using deep learning (DL). It requires large-scale training datasets and high computational and storage resources. Moreover, DL techniques and machine learning (ML) approaches in general, hypothesize that training and testing data come from the same domain, with the same input feature space and data distribution characteristics. This assumption, however, is not applicable in some real-world artificial intelligence (AI) applications. Moreover, there are situations where gathering real data is challenging, expensive, or rarely occurring, which can not meet the data requirements of DL models. deep transfer learning (DTL) has been introduced to overcome these issues, which helps develop high-performing models using real datasets that are small or slightly different but related to the training data. This paper presents a comprehensive survey of DTL-based ASR frameworks to shed light on the latest developments and helps academics and professionals understand current challenges. Specifically, after presenting the DTL background, a well-designed taxonomy is adopted to inform the state-of-the-art. A critical analysis is then conducted to identify the limitations and advantages of each framework. Moving on, a comparative study is introduced to highlight the current challenges before deriving opportunities for future research.
References in corpus (13)
- WaveNet: A Generative Model for Raw Audio
- Recent Advances in Domain Adaptation for the Classification of Remote Sensing Data
- A Survey on Transfer Learning in Natural Language Processing
- Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- A Study of the Generalizability of Self-Supervised Representations
- Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT
- Dropout Reduces Underfitting
- Long-span language modeling for speech recognition
- Semi-supervised acoustic modelling for five-lingual code-switched ASR using automatically-segmented soap opera speech
- Transfer learning from High-Resource to Low-Resource Language Improves Speech Affect Recognition Classification Accuracy
- The CUHK-TUDELFT System for The SLT 2021 Children Speech Recognition Challenge
- Multi-Modal Emotion Detection with Transfer Learning
Cited by in corpus (20)
- Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
- A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- Deep transfer learning for intrusion detection in industrial control networks: A comprehensive review
- Deep Learning for Steganalysis of Diverse Data Types: A review of methods, taxonomy, challenges and future directions
- Deep Reinforcement Learning for Intrusion Detection in IoT: A Survey
- Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives
- Advanced Deep Learning and Large Language Models: Comprehensive Insights for Cancer Detection
- Automatic Speech Recognition with BERT and CTC Transformers: A Review
- A Review of Deep Learning Approaches for Non-Invasive Cognitive Impairment Detection
- Machine Learning and Transformers for Thyroid Carcinoma Diagnosis: A Review
- Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications
- Hybrid Artificial Intelligence Strategies for Drone Navigation
- BreathAI: Transfer Learning-Based Thermal Imaging for Automated Breathing Pattern Recognition
- Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
- Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
- Transfer Learning-Based Deep Residual Learning for Speech Recognition in Clean and Noisy Environments
- Enhancing Cochlear Implant Signal Coding with Scaled Dot-Product Attention
- SE-Enhanced ViT and BiLSTM-Based Intrusion Detection for Secure IIoT and IoMT Environments
- Hybrid ResNet-1D-BiGRU with Multi-Head Attention for Cyberattack Detection in Industrial IoT Environments