GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
arXiv:2106.06909 · doi:10.21437/Interspeech.2021-1965
Abstract
This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality. Baseline systems are provided for popular speech recognition toolkits, namely Athena, ESPnet, Kaldi and Pika.
Cited by in corpus (27)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Self-Supervised Speech Representation Learning: A Review
- Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
- SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
- VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
- TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
- Acoustic correlates of the syllabic rhythm of speech: Modulation spectrum or local features of the temporal envelope
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
- Benchmarking Representations for Speech, Music, and Acoustic Events
- SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
- Complex Mapping between Neural Response Frequency and Linguistic Units in Natural Speech
- Advanced Long-Content Speech Recognition With Factorized Neural Transducer
- Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
- Domain Adaptation of low-resource Target-Domain models using well-trained ASR Conformer Models
- A Survey on Speech Large Language Models for Understanding
- Integrating Pause Information with Word Embeddings in Language Models for Alzheimer's Disease Detection from Spontaneous Speech
- WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
- SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
- AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
- Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech
- Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
- From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
- Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
- CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
- Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use