activity
20202022
most citedCitrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition

43 citations · 56 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS20224 cited

Multi-scale Speaker Diarization with Dynamic Scale Weighting

Tae Jin Park, Nithin Rao Koluguri, Jagadeesh Balam +1

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolutio…

eess.AS2021

CarneliNet: Neural Mixture Model for Automatic Speech Recognition

Aleksei Kalinov, Somshubra Majumdar, Jagadeesh Balam +1

End-to-end automatic speech recognition systems have achieved great accuracy by using deeper and deeper models. However, the increased depth comes with a larger receptive field tha…

cs.CL20219 cited

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar +10

In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitaliza…

eess.AS202143 cited

Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition

Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk +3

We propose Citrinet - a new end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. Citrinet is deep residual neural mo…

eess.AS2020

Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition

Jagadeesh Balam, Jocelyn Huang, Vitaly Lavrukhin +3

We present our experiments in training robust to noise an end-to-end automatic speech recognition (ASR) model using intensive data augmentation. We explore the efficacy of fine-tun…