WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
arXiv:2110.13900 · doi:10.1109/JSTSP.2022.3188113
Abstract
Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, spoken content, etc., learning universal representations for all speech tasks is challenging. To tackle the problem, we propose a new pre-trained model, WavLM, to solve full-stack downstream speech tasks. WavLM jointly learns masked speech prediction and denoising in pre-training. By this means, WavLM does not only keep the speech content modeling capability by the masked speech prediction, but also improves the potential to non-ASR tasks by the speech denoising. In addition, WavLM employs gated relative position bias for the Transformer structure to better capture the sequence ordering of input speech. We also scale up the training dataset from 60k hours to 94k hours. WavLM Large achieves state-of-the-art performance on the SUPERB benchmark, and brings significant improvements for various speech processing tasks on their representative benchmarks. The code and pre-trained models are available at https://aka.ms/wavlm.
Submitted to the Journal of Selected Topics in Signal Processing (JSTSP)
References in corpus (21)
- CogView: Mastering Text-to-Image Generation via Transformers
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Reducing Transformer Depth on Demand with Structured Dropout
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- Encoder-Decoder Based Attractors for End-to-End Neural Diarization
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- SUPERB: Speech processing Universal PERformance Benchmark
- Neural Speaker Diarization with Speaker-Wise Chain Rule
- The SpeakIn System for VoxCeleb Speaker Recognition Challange 2021
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
- Ultra Fast Speech Separation Model with Teacher Student Learning
- End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors
- UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
- Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech
- Have best of both worlds: two-pass hybrid and E2E cascading framework for speech recognition
- Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors
Cited by in corpus (111)
- Self-Supervised Speech Representation Learning: A Review
- HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
- BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
- VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
- Self-supervised representations in speech-based depression detection
- SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation
- The DiffuseStyleGesture+ entry to the GENEA Challenge 2023
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
- Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge
- PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
- Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
- C3-DINO: Joint Contrastive and Non-contrastive Self-Supervised Learning for Speaker Verification
- The VoxCeleb Speaker Recognition Challenge: A Retrospective
- Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
- PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
- DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
- Retrieval-Augmented Audio Deepfake Detection
- EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
- Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
- Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
- Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
- Speech foundation models on intelligibility prediction for hearing-impaired listeners
- Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
- Coding Speech through Vocal Tract Kinematics
- Benchmarking Representations for Speech, Music, and Acoustic Events
- Toward Improving Synthetic Audio Spoofing Detection Robustness via Meta-Learning and Disentangled Training With Adversarial Examples
- Training speaker recognition systems with limited data
- Distribution-based Emotion Recognition in Conversation
- Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
- Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
- Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
- Spectrogram features for audio and speech analysis
- A Backdoor Approach with Inverted Labels Using Dirty Label-Flipping Attacks
- Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
- Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation
- A Survey on Speech Large Language Models for Understanding
- PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
- Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
- Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
- Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
- Diffusion-Based Adversarial Purification for Speaker Verification
- An empirical study on speech restoration guided by self supervised speech representation
- Generic Speech Enhancement with Self-Supervised Representation Space Loss
- Audio-Language Datasets of Scenes and Events: A Survey
- Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
- Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
- Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
- Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
- On convex decision regions in deep network representations
- DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
- Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
- MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
- FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
- WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
- SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
- Phone and speaker spatial organization in self-supervised speech representations
- Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
- Convexity-based Pruning of Speech Representation Models
- Audio-visual child-adult speaker classification in dyadic interactions
- TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
- Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
- Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
- Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
- Semantics-Aware Human Motion Generation from Audio Instructions
- Beyond Neural-on-Neural Approaches to Speaker Gender Protection
- UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
- Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
- Deep Spectro-temporal Artifacts for Detecting Synthesized Speech
- Reference-free automatic speech severity evaluation using acoustic unit language modelling
- Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
- An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
- AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
- Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
- Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
- Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
- From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
- Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech
- A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
- DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
- Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
- Multi-Resolution Generative Modeling of Human Motion from Limited Data
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
- Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
- Multilingual Stutter Event Detection for English, German, and Mandarin Speech
- A survey of AI-generated voices and their detection
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Speaker Group Encoding in Self-supervised Speech Recognition Models
- On feature representations for marmoset vocal communication analysis
- Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
- Investigating self-supervised features for expressive, multilingual voice conversion
- CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
- Test-Time Adaptation For Speech Enhancement Via Mask Polarization
- FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
- SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
- Transformer-based Automatic Speech Recognition of Formal and Colloquial Czech in MALACH Project
- MMMOS: Multi-domain Multi-axis Audio Quality Assessment
- Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
- An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
- Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
- Semantic enrichment towards efficient speech representations
- Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024