43 citations · 56 across the 4 of their papers we have counts for
5 papers
Multi-scale Speaker Diarization with Dynamic Scale Weighting
Tae Jin Park, Nithin Rao Koluguri, Jagadeesh Balam +1
Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolutio…
CarneliNet: Neural Mixture Model for Automatic Speech Recognition
Aleksei Kalinov, Somshubra Majumdar, Jagadeesh Balam +1
End-to-end automatic speech recognition systems have achieved great accuracy by using deeper and deeper models. However, the increased depth comes with a larger receptive field tha…
SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar +10
In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitaliza…
Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk +3
We propose Citrinet - a new end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. Citrinet is deep residual neural mo…
Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition
Jagadeesh Balam, Jocelyn Huang, Vitaly Lavrukhin +3
We present our experiments in training robust to noise an end-to-end automatic speech recognition (ASR) model using intensive data augmentation. We explore the efficacy of fine-tun…