TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
arXiv:2007.06028 · doi:10.1109/TASLP.2021.3095662
Abstract
We introduce a self-supervised speech pre-training method called TERA, which stands for Transformer Encoder Representations from Alteration. Recent approaches often learn by using a single auxiliary task like contrastive prediction, autoregressive prediction, or masked reconstruction. Unlike previous methods, we use alteration along three orthogonal axes to pre-train Transformer Encoders on a large amount of unlabeled speech. The model learns through the reconstruction of acoustic frames from their altered counterpart, where we use a stochastic policy to alter along various dimensions: time, frequency, and magnitude. TERA can be used for speech representations extraction or fine-tuning with downstream models. We evaluate TERA on several downstream tasks, including phoneme classification, keyword spotting, speaker recognition, and speech recognition. We present a large-scale comparison of various self-supervised models. TERA achieves strong performance in the comparison by improving upon surface features and outperforming previous models. In our experiments, we study the effect of applying different alteration techniques, pre-training on more data, and pre-training on various features. We analyze different model sizes and find that smaller models are strong representation learners than larger models, while larger models are more effective for downstream fine-tuning than smaller models. Furthermore, we show the proposed method is transferable to downstream datasets not used in pre-training.
Published in IEEE/ACM TASLP, final published article available at https://ieeexplore.ieee.org/document/9478264
References in corpus (6)
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Self-supervised audio representation learning for mobile devices
- Learning audio representations via phase prediction
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- Masked Pre-trained Encoder base on Joint CTC-Transformer
Cited by in corpus (61)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Self-Supervised Speech Representation Learning: A Review
- Transfer learning based physics-informed neural networks for solving inverse problems in engineering structures under different loading scenarios
- Keyword Transformer: A Self-Attention Model for Keyword Spotting
- Nearest Neighbor-Based Contrastive Learning for Hyperspectral and LiDAR Data Classification
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- COCOA: Cross Modality Contrastive Learning for Sensor Data
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech
- SUPERB: Speech processing Universal PERformance Benchmark
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- ATST: Audio Representation Learning with Teacher-Student Transformer
- Semi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On Federated Learning using Multiview Pseudo-Labeling
- A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
- Non-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech
- Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- DECAR: Deep Clustering for learning general-purpose Audio Representations
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
- Adversarial defense for automatic speaker verification by cascaded self-supervised learning models
- Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition
- SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
- RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing
- Towards Semi-Supervised Semantics Understanding from Speech
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation
- Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation
- Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning
- Style Attuned Pre-training and Parameter Efficient Fine-tuning for Spoken Language Understanding
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
- End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Self-supervised Speaker Recognition Training Using Human-Machine Dialogues
- Delta Keyword Transformer: Bringing Transformers to the Edge through Dynamically Pruned Multi-Head Self-Attention
- Audio MFCC-gram Transformers for respiratory insufficiency detection in COVID-19
- Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
- Complementing Handcrafted Features with Raw Waveform Using a Light-weight Auxiliary Model
- Convexity-based Pruning of Speech Representation Models
- Similarity Analysis of Self-Supervised Speech Representations
- Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
- Fusion of Embeddings Networks for Robust Combination of Text Dependent and Independent Speaker Recognition
- Using Self-Supervised Feature Extractors with Attention for Automatic COVID-19 Detection from Speech
- SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
- Representation Learning for Sequence Data with Deep Autoencoding Predictive Components
- Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation
- Transformer-based Automatic Speech Recognition of Formal and Colloquial Czech in MALACH Project
- SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
- Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training
- Meta learning to classify intent and slot labels with noisy few shot examples
- Semantic enrichment towards efficient speech representations
- Attention-Free Keyword Spotting
- Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency