THCHS-30 : A Free Chinese Speech Corpus
arXiv:1512.01882
Abstract
Speech data is crucially important for speech recognition research. There are quite some speech databases that can be purchased at prices that are reasonable for most research institutes. However, for young people who just start research activities or those who just gain initial interest in this direction, the cost for data is still an annoying barrier. We support the `free data' movement in speech recognition: research institutes (particularly supported by public funds) publish their data freely so that new researchers can obtain sufficient data to kick of their career. In this paper, we follow this trend and release a free Chinese speech database THCHS-30 that can be used to build a full- edged Chinese speech recognition system. We report the baseline system established with this database, including the performance under highly noisy conditions.
References in corpus (1)
Cited by in corpus (32)
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- ATCSpeech: a multilingual pilot-controller speech corpus from real Air Traffic Control environment
- USTC-NELSLIP System Description for DIHARD-III Challenge
- The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units
- BERT-LID: Leveraging BERT to Improve Spoken Language Identification
- Investigation of Using Disentangled and Interpretable Representations for One-shot Cross-lingual Voice Conversion
- Universal Phone Recognition with a Multilingual Allophone System
- Residual acoustic echo suppression based on efficient multi-task convolutional neural network
- Gated Recurrent Unit Based Acoustic Modeling with Future Context
- A Novel Speech-Driven Lip-Sync Model with CNN and LSTM
- AP17-OLR Challenge: Data, Plan, and Baseline
- DeepSinger: Singing Voice Synthesis with Data Mined From the Web
- Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition
- OC16-CE80: A Chinese-English Mixlingual Database and A Speech Recognition Baseline
- DiDiSpeech: A Large Scale Mandarin Speech Corpus
- Wideband Audio Waveform Evaluation Networks: Efficient, Accurate Estimation of Speech Qualities
- a novel cross-lingual voice cloning approach with a few text-free samples
- Acoustic Word Embedding System for Code-Switching Query-by-example Spoken Term Detection
- AP20-OLR Challenge: Three Tasks and Their Baselines
- AlloVera: A Multilingual Allophone Database
- MLNET: An Adaptive Multiple Receptive-field Attention Neural Network for Voice Activity Detection
- AP18-OLR Challenge: Three Tasks and Their Baselines
- Improving Gated Recurrent Unit Based Acoustic Modeling with Batch Normalization and Enlarged Context
- Word-Free Spoken Language Understanding for Mandarin-Chinese
- AP19-OLR Challenge: Three Tasks and Their Baselines
- Additive Phoneme-aware Margin Softmax Loss for Language Recognition
- Empowering cyberphysical systems of systems with intelligence
- How Far Are We from Robust Voice Conversion: A Survey
- Oriental Language Recognition (OLR) 2020: Summary and Analysis
- Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement
- Exploring Teacher-Student Learning Approach for Multi-lingual Speech-to-Intent Classification