activity
20172021
most citedAccent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

11 citations · 29 across the 8 of their papers we have counts for

collaborators

10 papers

eess.AS202111 cited

Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

Jialu Li, Vimal Manohar, Pooja Chitkara +5

Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…

eess.AS20211 cited

Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution

Yangyang Shi, Chunyang Wu, Dilin Wang +9

This paper improves the streaming transformer transducer for speech recognition by using non-causal convolution. Many works apply the causal convolution to improve streaming transf…

eess.AS20211 cited

On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models

Xiaohui Zhang, Vimal Manohar, David Zhang +7

Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…

eess.AS2020

Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR

Xiaohui Zhang, Frank Zhang, Chunxi Liu +8

In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on thre…

eess.AS20202 cited

Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces

Frank Zhang, Yongqiang Wang, Xiaohui Zhang +3

In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state…

eess.AS20199 cited

From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition

Duc Le, Xiaohui Zhang, Weiyi Zheng +3

There is an implicit assumption that traditional hybrid approaches for automatic speech recognition (ASR) cannot directly model graphemes and need to rely on phonetic lexicons to g…