Decoupling recognition and transcription in Mandarin ASR
arXiv:2108.01129
Abstract
Much of the recent literature on automatic speech recognition (ASR) is taking an end-to-end approach. Unlike English where the writing system is closely related to sound, Chinese characters (Hanzi) represent meaning, not sound. We propose factoring audio -> Hanzi into two sub-tasks: (1) audio -> Pinyin and (2) Pinyin -> Hanzi, where Pinyin is a system of phonetic transcription of standard Chinese. Factoring the audio -> Hanzi task in this way achieves 3.9% CER (character error rate) on the Aishell-1 corpus, the best result reported on this dataset so far.
submitted to ASRU 2021
References in corpus (21)
- Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- Recent Developments on ESPnet Toolkit Boosted by Conformer
- Multi-head Monotonic Chunkwise Attention For Online Speech Recognition
- Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- Improved Conformer-based End-to-End Speech Recognition Using Neural Architecture Search
- Intermediate Loss Regularization for CTC-based Speech Recognition
- Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
- Improving RNN transducer with normalized jointer network
- Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition
- Decoupling Pronunciation and Language for End-to-end Code-switching Automatic Speech Recognition
- Transformer with Bidirectional Decoder for Speech Recognition
- Gated Recurrent Fusion with Joint Training Framework for Robust End-to-End Speech Recognition
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-based LVCSR
- Automatic recognition of suprasegmentals in speech
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- Head-synchronous Decoding for Transformer-based Streaming ASR