activity
20192021
most citedAccent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

11 citations · 27 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS202111 cited

Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

Jialu Li, Vimal Manohar, Pooja Chitkara +5

Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…

eess.AS20211 cited

On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models

Xiaohui Zhang, Vimal Manohar, David Zhang +7

Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…

eess.AS2020

Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR

Xiaohui Zhang, Frank Zhang, Chunxi Liu +8

In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on thre…

eess.AS20205 cited

Weak-Attention Suppression For Transformer Based Speech Recognition

Yangyang Shi, Yongqiang Wang, Chunyang Wu +5

Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acousti…

eess.AS20201 cited

Streaming Transformer-based Acoustic Models Using Self-attention with Augmented Memory

Chunyang Wu, Yongqiang Wang, Yangyang Shi +2

Transformer-based acoustic modeling has achieved great suc-cess for both hybrid and sequence-to-sequence speech recogni-tion. However, it requires access to the full sequence, and…

eess.AS20202 cited

Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces

Frank Zhang, Yongqiang Wang, Xiaohui Zhang +3

In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state…