Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
arXiv:2005.08575
Abstract
For self-supervised speech processing, it is crucial to use pretrained models as speech representation extractors. In recent works, increasing the size of the model has been utilized in acoustic model training in order to achieve better performance. In this paper, we propose Audio ALBERT, a lite version of the self-supervised speech representation model. We use the representations with two downstream tasks, speaker identification, and phoneme classification. We show that Audio ALBERT is capable of achieving competitive performance with those huge models in the downstream tasks while utilizing 91\% fewer parameters. Moreover, we use some simple probing models to measure how much the information of the speaker and phoneme is encoded in latent representations. In probing experiments, we find that the latent representations encode richer information of both phoneme and speaker than that of the last layer.
Accepted by IEEE Spoken Language Technology Workshop 2021
References in corpus (6)
- A Simple Framework for Contrastive Learning of Visual Representations
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks
Cited by in corpus (15)
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- On the use of Self-supervised Pre-trained Acoustic and Linguistic Features for Continuous Speech Emotion Recognition
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- One4all User Representation for Recommender Systems in E-commerce
- What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure
- Towards Semi-Supervised Semantics Understanding from Speech
- Style Attuned Pre-training and Parameter Efficient Fine-tuning for Spoken Language Understanding
- Probing Acoustic Representations for Phonetic Properties
- Translate Reverberated Speech to Anechoic Ones: Speech Dereverberation with BERT
- Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
- Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency
- Using Pause Information for More Accurate Entity Recognition
- A Transformer Based Pitch Sequence Autoencoder with MIDI Augmentation
- Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers