Libri-Light: A Benchmark for ASR with Limited or No Supervision
arXiv:1912.07875 · doi:10.1109/ICASSP40776.2020.9052942
Abstract
We introduce a new collection of spoken English audio suitable for training speech recognition systems under limited or no supervision. It is derived from open-source audio books from the LibriVox project. It contains over 60K hours of audio, which is, to our knowledge, the largest freely-available corpus of speech. The audio has been segmented using voice activity detection and is tagged with SNR, speaker ID and genre descriptions. Additionally, we provide baseline systems and evaluation metrics working under three settings: (1) the zero resource/unsupervised setting (ABX), (2) the semi-supervised setting (PER, CER) and (3) the distant supervision setting (WER). Settings (2) and (3) use limited textual resources (10 minutes to 10 hours) aligned with the speech. Setting (3) uses large amounts of unaligned text. They are evaluated on the standard LibriSpeech dev and test sets for comparison with the supervised state-of-the-art.
References in corpus (10)
- Representation Learning with Contrastive Predictive Coding
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Self-Training for End-to-End Speech Recognition
- wav2letter++: The Fastest Open-source Speech Recognition System
- RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation
- Effectiveness of self-supervised pre-training for speech recognition
- Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
- Letter-Based Speech Recognition with Gated ConvNets
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
- Almost-unsupervised Speech Recognition with Close-to-zero Resource Based on Phonetic Structures Learned from Very Small Unpaired Speech and Text Data
Cited by in corpus (79)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- SpeechBrain: A General-Purpose Speech Toolkit
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- Self-Supervised Speech Representation Learning: A Review
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- Improved Noisy Student Training for Automatic Speech Recognition
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- Self-supervised Pretraining of Visual Features in the Wild
- A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion
- Effectiveness of self-supervised pre-training for speech recognition
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- Data Augmenting Contrastive Learning of Speech Representations in the Time Domain
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- Unsupervised Automatic Speech Recognition: A Review
- Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
- Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers
- On the use of Self-supervised Pre-trained Acoustic and Linguistic Features for Continuous Speech Emotion Recognition
- Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- TRILLsson: Distilled Universal Paralinguistic Speech Representations
- Universal Paralinguistic Speech Representations Using Self-Supervised Conformers
- Multi-Format Contrastive Learning of Audio Representations
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
- Semi-Supervised Speech Recognition via Local Prior Matching
- Iterative Pseudo-Labeling for Speech Recognition
- Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge
- From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
- Look Once to Hear: Target Speech Hearing with Noisy Examples
- Analysing Discrete Self Supervised Speech Representation for Spoken Language Modeling
- Word Segmentation on Discovered Phone Units with Dynamic Programming and Self-Supervised Scoring
- Unsupervised Speech Recognition
- A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
- Pushing the Limits of Non-Autoregressive Speech Recognition
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- Ultra Fast Speech Separation Model with Teacher Student Learning
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- Self-training and Pre-training are Complementary for Speech Recognition
- BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
- Learning Robust and Multilingual Speech Representations
- Magic dust for cross-lingual adaptation of monolingual wav2vec-2.0
- Variable-rate discrete representation learning
- Lhotse: a speech data representation library for the modern deep learning ecosystem
- Benchmarking Representations for Speech, Music, and Acoustic Events
- SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
- Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
- Large-Scale Self- and Semi-Supervised Learning for Speech Translation
- Hi-Fi Multi-Speaker English TTS Dataset
- End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations
- Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
- Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware Modeling
- Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement
- Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
- Self-Training for End-to-End Speech Translation
- Residual Energy-Based Models for End-to-End Speech Recognition
- Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
- WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
- Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
- Guided contrastive self-supervised pre-training for automatic speech recognition
- Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
- Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech
- BSTC: A Large-Scale Chinese-English Speech Translation Dataset
- The effectiveness of unsupervised subword modeling with autoregressive and cross-lingual phone-aware networks
- Improving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Data
- Similarity Analysis of Self-Supervised Speech Representations
- Fast Development of ASR in African Languages using Self Supervised Speech Representation Learning
- Reference-free automatic speech severity evaluation using acoustic unit language modelling
- CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
- Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
- XPPG-PCA: Reference-free automatic speech severity evaluation with principal components
- Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?
- Topic Model Robustness to Automatic Speech Recognition Errors in Podcast Transcripts
- Empowering cyberphysical systems of systems with intelligence
- Learning When to Trust Which Teacher for Weakly Supervised ASR