activity
20182023
most citedTransformer-Transducer: End-to-End Speech Recognition with Self-Attention

66 citations · 125 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2022

SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning

Tzu-hsun Feng, Annie Dong, Ching-Feng Yeh +11

We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge buil…

cs.CL2021

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

Yangyang Shi, Varun Nagaraja, Chunyang Wu +9

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…

cs.CL2021★ 5 cited

Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding

Suyoun Kim, Abhinav Arora, Duc Le +4

Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator…

cs.CL2020

Alignment Restricted Streaming Recurrent Neural Network Transducer

Jay Mahadeokar, Yuan Shangguan, Duc Le +6

There is a growing interest in the speech community in developing Recurrent Neural Network Transducer (RNN-T) models for automatic speech recognition (ASR) applications. RNN-T is t…

cs.CL2020

Streaming Attention-Based Models with Augmented Memory for End-to-End Speech Recognition

Ching-Feng Yeh, Yongqiang Wang, Yangyang Shi +4

Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation and automatic speech recognition. One m…

cs.CL2020★ 1 cited

Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications

Yongqiang Wang, Yangyang Shi, Frank Zhang +4

In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications. We compare the…