Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
arXiv:1910.12638 · doi:10.1109/ICASSP40776.2020.9054458
Abstract
We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech. Previous speech representation methods learn through conditioning on past frames and predicting information about future frames. Whereas Mockingjay is designed to predict the current frame through jointly conditioning on both past and future contexts. The Mockingjay representation improves performance for a wide range of downstream tasks, including phoneme classification, speaker recognition, and sentiment classification on spoken content, while outperforming other approaches. Mockingjay is empirically powerful and can be fine-tuned with downstream models, with only 2 epochs we further improve performance dramatically. In a low resource setting with only 0.1% of labeled data, we outperform the result of Mel-features that uses all 100% labeled data.
Accepted by ICASSP 2020, Lecture Session
References in corpus (9)
- Adam: A Method for Stochastic Optimization
- Representation Learning with Contrastive Predictive Coding
- Convolutional Sequence to Sequence Learning
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Layer Normalization
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Unsupervised speech representation learning using WaveNet autoencoders
- Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
- Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder
Cited by in corpus (96)
- On the Opportunities and Risks of Foundation Models
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Self-Supervised Speech Representation Learning: A Review
- Dawn of the transformer era in speech emotion recognition: closing the valence gap
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- StorSeismic: A new paradigm in deep learning for seismic processing
- BYOL for Audio: Exploring Pre-trained General-purpose Audio Representations
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- On the use of Self-supervised Pre-trained Acoustic and Linguistic Features for Continuous Speech Emotion Recognition
- SUPERB: Speech processing Universal PERformance Benchmark
- Self-Supervised learning with cross-modal transformers for emotion recognition
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
- Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
- BERTphone: Phonetically-Aware Encoder Representations for Utterance-Level Speaker and Language Recognition
- ATST: Audio Representation Learning with Teacher-Student Transformer
- C3-DINO: Joint Contrastive and Non-contrastive Self-Supervised Learning for Speaker Verification
- Semi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On Federated Learning using Multiview Pseudo-Labeling
- A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
- Non-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech
- Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition
- FRILL: A Non-Semantic Speech Embedding for Mobile Devices
- Self-Supervised Learning of Audio Representations from Permutations with Differentiable Ranking
- Contrastive Learning with Positive-Negative Frame Mask for Music Representation
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- Representation Selective Self-distillation and wav2vec 2.0 Feature Exploration for Spoof-aware Speaker Verification
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- DECAR: Deep Clustering for learning general-purpose Audio Representations
- MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
- Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks
- Defense for Black-box Attacks on Anti-spoofing Models by Self-Supervised Learning
- CSTNet: Contrastive Speech Translation Network for Self-Supervised Speech Representation Learning
- UniSpeech at scale: An Empirical Study of Pre-training Method on Large-Scale Speech Recognition Dataset
- Adversarial defense for automatic speaker verification by cascaded self-supervised learning models
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning
- Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models
- What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure
- Wave to Syntax: Probing spoken language models for syntax
- Vector-Quantized Autoregressive Predictive Coding
- Masked Pre-trained Encoder base on Joint CTC-Transformer
- Towards Semi-Supervised Semantics Understanding from Speech
- Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- Conditional independence for pretext task selection in Self-supervised speech representation learning
- Saturn Platform: Foundation Model Operations and Generative AI for Financial Services
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Self-supervised Speaker Recognition Training Using Human-Machine Dialogues
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
- Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning
- Audio MFCC-gram Transformers for respiratory insufficiency detection in COVID-19
- Understanding Self-Attention of Self-Supervised Audio Transformers
- Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining
- Generic Speech Enhancement with Self-Supervised Representation Space Loss
- DiDiSpeech: A Large Scale Mandarin Speech Corpus
- Adversarially learning disentangled speech representations for robust multi-factor voice conversion
- Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
- Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning
- Probing Acoustic Representations for Phonetic Properties
- S2VC: A Framework for Any-to-Any Voice Conversion with Self-Supervised Pretrained Representations
- Contrastive Semi-supervised Learning for ASR
- Phone and speaker spatial organization in self-supervised speech representations
- Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese
- Guided contrastive self-supervised pre-training for automatic speech recognition
- Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- Translate Reverberated Speech to Anechoic Ones: Speech Dereverberation with BERT
- Fusion of Embeddings Networks for Robust Combination of Text Dependent and Independent Speaker Recognition
- Using Self-Supervised Feature Extractors with Attention for Automatic COVID-19 Detection from Speech
- Any-to-One Sequence-to-Sequence Voice Conversion using Self-Supervised Discrete Speech Representations
- Unsupervised Representation Learning for Speaker Recognition via Contrastive Equilibrium Learning
- SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
- Similarity Analysis of Self-Supervised Speech Representations
- Improved Speech Representations with Multi-Target Autoregressive Predictive Coding
- Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-training and Its Application to Children's ASR
- Phoneme-based Distribution Regularization for Speech Enhancement
- Using Pause Information for More Accurate Entity Recognition
- Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers
- Pretext Tasks selection for multitask self-supervised speech representation learning
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
- Auto-KWS 2021 Challenge: Task, Datasets, and Baselines
- Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency
- Semantic enrichment towards efficient speech representations
- Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation
- Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering
- Towards Language Modelling in the Speech Domain Using Sub-word Linguistic Units
- BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks
- Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?
- Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training