66 citations · 156 across the 21 of their papers we have counts for
17 papers · 1 filter
Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
Andros Tjandra, Nayan Singhal, David Zhang +4
End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high…
Streaming parallel transducer beam search with fast-slow cascaded encoders
Jay Mahadeokar, Yangyang Shi, Ke Li +5
Streaming ASR with strict latency constraints is required in many speech recognition applications. In order to achieve the required latency, streaming ASR models sacrifice accuracy…
Neural-FST Class Language Model for End-to-End Speech Recognition
Antoine Bruguier, Duc Le, Rohit Prabhavalkar +7
We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transduce…
Collaborative Training of Acoustic Encoders for Speech Recognition
Varun Nagaraja, Yangyang Shi, Ganesh Venkatesh +3
On-device speech recognition requires training models of different sizes for deploying on devices with various computational budgets. When building such different models, we can be…
Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
Yangyang Shi, Varun Nagaraja, Chunyang Wu +9
We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…
Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
Suyoun Kim, Abhinav Arora, Duc Le +4
Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator…