activity
20192024
most citedTransformer-Transducer: End-to-End Speech Recognition with Self-Attention

66 citations · 141 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL20221 cited

Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities

Andros Tjandra, Nayan Singhal, David Zhang +4

End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high…

cs.CL2022

Joint Audio/Text Training for Transformer Rescorer of Streaming Speech Recognition

Suyoun Kim, Ke Li, Lucas Kabela +4

Recently, there has been an increasing interest in two-pass streaming end-to-end speech recognition (ASR) that incorporates a 2nd-pass rescoring model on top of the conventional 1s…

cs.CL2022

Streaming parallel transducer beam search with fast-slow cascaded encoders

Jay Mahadeokar, Yangyang Shi, Ke Li +5

Streaming ASR with strict latency constraints is required in many speech recognition applications. In order to achieve the required latency, streaming ASR models sacrifice accuracy…

cs.CL2022

Neural-FST Class Language Model for End-to-End Speech Recognition

Antoine Bruguier, Duc Le, Rohit Prabhavalkar +7

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transduce…

cs.CL2021

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

Yangyang Shi, Varun Nagaraja, Chunyang Wu +9

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…

cs.CL20215 cited

Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding

Suyoun Kim, Abhinav Arora, Duc Le +4

Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator…