activity
20192021
most citedAccent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

11 citations · 27 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20201 cited

Improving RNN Transducer Based ASR with Auxiliary Tasks

Chunxi Liu, Frank Zhang, Duc Le +3

End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recogni…

cs.CL2020

Streaming Attention-Based Models with Augmented Memory for End-to-End Speech Recognition

Ching-Feng Yeh, Yongqiang Wang, Yangyang Shi +4

Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation and automatic speech recognition. One m…

cs.CL20201 cited

Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications

Yongqiang Wang, Yangyang Shi, Frank Zhang +4

In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications. We compare the…

cs.CL20205 cited

Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model

Da-Rong Liu, Chunxi Liu, Frank Zhang +3

Videos uploaded on social media are often accompanied with textual descriptions. In building automatic speech recognition (ASR) systems for videos, we can exploit the contextual in…

cs.CL2019

Training ASR models by Generation of Contextual Information

Kritika Singh, Dmytro Okhonko, Jun Liu +8

Supervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales,…

cs.CL2019

Deja-vu: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

Andros Tjandra, Chunxi Liu, Frank Zhang +5

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to f…