activity
20182021
most citedTransformer-Transducer: End-to-End Speech Recognition with Self-Attention

66 citations · 103 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2021

Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion

Duc Le, Mahaveer Jain, Gil Keren +9

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for sp…

eess.AS2020

Contextual RNN-T For Open Domain ASR

Mahaveer Jain, Gil Keren, Jay Mahadeokar +3

End-to-end (E2E) systems for automatic speech recognition (ASR), such as RNN Transducer (RNN-T) and Listen-Attend-Spell (LAS) blend the individual components of a traditional hybri…

cs.CL20193 cited

AIPNet: Generative Adversarial Pre-training of Accent-invariant Networks for End-to-end Speech Recognition

Yi-Chen Chen, Zhaojun Yang, Ching-Feng Yeh +2

As one of the major sources in speech variability, accents have posed a grand challenge to the robustness of speech recognition systems. In this paper, our goal is to build a unifi…

cs.CL201934 cited

RNN-T For Latency Controlled ASR With Improved Beam Search

Mahaveer Jain, Kjell Schubert, Jay Mahadeokar +5

Neural transducer-based systems such as RNN Transducers (RNN-T) for automatic speech recognition (ASR) blend the individual components of a traditional hybrid ASR systems (acoustic…

eess.AS201966 cited

Transformer-Transducer: End-to-End Speech Recognition with Self-Attention

Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar +6

We explore options to use Transformer networks in neural transducer for end-to-end speech recognition. Transformer networks use self-attention for sequence modeling and comes with…

cs.CL2018

End-to-end contextual speech recognition using class language models and a token passing decoder

Zhehuai Chen, Mahaveer Jain, Yongqiang Wang +2

End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a unified model. Although it simplifies tr…