vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
arXiv:1910.05453
Abstract
We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables the direct application of algorithms from the NLP community which require discrete inputs. Experiments show that BERT pre-training achieves a new state of the art on TIMIT phoneme classification and WSJ speech recognition.
References in corpus (4)
Cited by in corpus (19)
- Contrastive Representation Learning: A Framework and Review
- An Overview of Indian Spoken Language Recognition from Machine Learning Perspective
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Privacy-preserving Voice Analysis via Disentangled Representations
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- Adapting Multilingual Speech Representation Model for a New, Underresourced Language through Multilingual Fine-tuning and Continued Pretraining
- DiscreTalk: Text-to-Speech as a Machine Translation Problem
- Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals
- PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice Conversion
- Learning Robust and Multilingual Speech Representations
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector Quantization
- Vector Quantized Contrastive Predictive Coding for Template-based Music Generation
- Indonesian Automatic Speech Recognition with XLSR-53
- Towards Semi-Supervised Semantics Understanding from Speech
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
- Audio MFCC-gram Transformers for respiratory insufficiency detection in COVID-19
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Improving speech recognition models with small samples for air traffic control systems