11 citations · 28 across the 10 of their papers we have counts for
8 papers · 1 filter
Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
Jialu Li, Vimal Manohar, Pooja Chitkara +5
Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…
On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models
Xiaohui Zhang, Vimal Manohar, David Zhang +7
Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…
Kaizen: Continuously improving teacher using Exponential Moving Average for semi-supervised speech recognition
Vimal Manohar, Tatiana Likhomanenko, Qiantong Xu +5
In this paper, we introduce the Kaizen framework that uses a continuously improving teacher to generate pseudo-labels for semi-supervised speech recognition (ASR). The proposed app…
Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR
Xiaohui Zhang, Frank Zhang, Chunxi Liu +8
In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on thre…
Contextual RNN-T For Open Domain ASR
Mahaveer Jain, Gil Keren, Jay Mahadeokar +3
End-to-end (E2E) systems for automatic speech recognition (ASR), such as RNN Transducer (RNN-T) and Listen-Attend-Spell (LAS) blend the individual components of a traditional hybri…
Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
Frank Zhang, Yongqiang Wang, Xiaohui Zhang +3
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state…