activity
20152021
most citedLearning Latent Spatio-Temporal Compositional Model for Human Action Recognition

27 citations · 78 across the 10 of their papers we have counts for

collaborators

26 papers

eess.AS2021

Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction

David Qiu, Yanzhang He, Qiujia Li +3

Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utter…

cs.CL2021

Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models

Thibault Doutre, Wei Han, Chung-Cheng Chiu +3

Streaming end-to-end automatic speech recognition (ASR) systems are widely used in everyday applications that require transcribing speech to text in real-time. Their minimal latenc…

eess.AS20211 cited

Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models

Zhiyun Lu, Wei Han, Yu Zhang +1

Although end-to-end automatic speech recognition (e2e ASR) models are widely deployed in many applications, there have been very few studies to understand models' robustness agains…

eess.AS20212 cited

Learning Word-Level Confidence For Subword End-to-End ASR

David Qiu, Qiujia Li, Yanzhang He +9

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed trainin…

eess.AS2021

Residual Energy-Based Models for End-to-End Speech Recognition

Qiujia Li, Yu Zhang, Bo Li +2

End-to-end models with auto-regressive decoders have shown impressive results for automatic speech recognition (ASR). These models formulate the sequence-level probability as a pro…

cs.CV20203 cited

Spatial-Temporal Alignment Network for Action Recognition and Detection

Junwei Liang, Liangliang Cao, Xuehan Xiong +2

This paper studies how to introduce viewpoint-invariant feature representations that can help action recognition and detection. Although we have witnessed great progress of action…