Self-Supervised Speech Representation Learning: A Review
arXiv:2205.10643 · doi:10.1109/JSTSP.2022.3207050
Abstract
Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply this to dialects and languages for which only limited labeled data is available. Self-supervised representation learning methods promise a single universal model that would benefit a wide variety of tasks and domains. Such methods have shown success in natural language processing and computer vision domains, achieving new levels of performance while reducing the number of labels required for many downstream scenarios. Speech representation learning is experiencing similar progress in three main categories: generative, contrastive, and predictive methods. Other approaches rely on multi-modal data for pre-training, mixing text or visual data streams with speech. Although self-supervised speech representation is still a nascent research area, it is closely related to acoustic word embedding and learning with zero lexical resources, both of which have seen active research for many years. This review presents approaches for self-supervised speech representation learning and their connection to other research areas. Since many current methods focus solely on automatic speech recognition as a downstream task, we review recent efforts on benchmarking learned representations to extend the application beyond speech recognition.
References in corpus (20)
- On the Opportunities and Risks of Foundation Models
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- On Using Monolingual Corpora in Neural Machine Translation
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Deep Variational Canonical Correlation Analysis
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Self-supervised audio representation learning for mobile devices
- DiscreTalk: Text-to-Speech as a Machine Translation Problem
- Are discrete units necessary for Spoken Language Modeling?
- Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem
- Learning audio representations via phase prediction
- Masked Pre-trained Encoder base on Joint CTC-Transformer
Cited by in corpus (20)
- Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study
- VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
- SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
- Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
- Distribution-based Emotion Recognition in Conversation
- SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
- Generic Speech Enhancement with Self-Supervised Representation Space Loss
- Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
- Experimenting with Additive Margins for Contrastive Self-Supervised Speaker Verification
- Convexity-based Pruning of Speech Representation Models
- UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
- Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
- AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
- Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
- RECA-PD: A Robust Explainable Cross-Attention Method for Speech-based Parkinson's Disease Classification
- Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
- Cross-Corpora Spoken Language Identification with Domain Diversification and Generalization