activity
20202022
most citedBest of Both Worlds: Robust Accented Speech Recognition with Adversarial Transfer Learning

2 citations · 2 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2022

Towards Personalization of CTC Speech Recognition Models with Contextual Adapters and Adaptive Boosting

Saket Dingliwal, Monica Sunkara, Sravan Bodapati +3

End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregr…

eess.AS2021

Remember the context! ASR slot error correction through memorization

Dhanush Bekal, Ashish Shenoy, Monica Sunkara +2

Accurate recognition of slot values such as domain specific words or named entities by automatic speech recognition (ASR) systems forms the core of the Goal-oriented Dialogue Syste…

cs.CL2021

Adapting Long Context NLM for ASR Rescoring in Conversational Agents

Ashish Shenoy, Sravan Bodapati, Monica Sunkara +2

Neural Language Models (NLM), when trained and evaluated with context spanning multiple utterances, have been shown to consistently outperform both conventional n-gram language mod…

eess.AS20212 cited

Best of Both Worlds: Robust Accented Speech Recognition with Adversarial Transfer Learning

Nilaksh Das, Sravan Bodapati, Monica Sunkara +2

Training deep neural networks for automatic speech recognition (ASR) requires large amounts of transcribed speech. This becomes a bottleneck for training robust models for accented…

cs.CL2021

Neural Inverse Text Normalization

Monica Sunkara, Chaitanya Shivade, Sravan Bodapati +1

While there have been several contributions exploring state of the art techniques for text normalization, the problem of inverse text normalization (ITN) remains relatively unexplo…

eess.AS2020

Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech

Monica Sunkara, Srikanth Ronanki, Dhanush Bekal +2

In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data.…