activity
20192021
most citedAccent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

11 citations · 27 across the 11 of their papers we have counts for

collaborators

14 papers

eess.AS202111 cited

Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

Jialu Li, Vimal Manohar, Pooja Chitkara +5

Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…

eess.AS20211 cited

On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models

Xiaohui Zhang, Vimal Manohar, David Zhang +7

Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…

eess.AS2020

Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR

Xiaohui Zhang, Frank Zhang, Chunxi Liu +8

In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on thre…

cs.CL20201 cited

Improving RNN Transducer Based ASR with Auxiliary Tasks

Chunxi Liu, Frank Zhang, Duc Le +3

End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recogni…

cs.CL2020

Streaming Attention-Based Models with Augmented Memory for End-to-End Speech Recognition

Ching-Feng Yeh, Yongqiang Wang, Yangyang Shi +4

Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation and automatic speech recognition. One m…

cs.CL20201 cited

Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications

Yongqiang Wang, Yangyang Shi, Frank Zhang +4

In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications. We compare the…