1 citations · 1 across the 4 of their papers we have counts for
7 papers · 1 filter
Device Directedness with Contextual Cues for Spoken Dialog Systems
Dhanush Bekal, Sundararajan Srinivasan, Sravan Bodapati +2
In this work, we define barge-in verification as a supervised learning task where audio-only information is used to classify user spoken dialogue into true and false barge-ins. Fol…
Towards Personalization of CTC Speech Recognition Models with Contextual Adapters and Adaptive Boosting
Saket Dingliwal, Monica Sunkara, Sravan Bodapati +3
End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregr…
Adapting Long Context NLM for ASR Rescoring in Conversational Agents
Ashish Shenoy, Sravan Bodapati, Monica Sunkara +2
Neural Language Models (NLM), when trained and evaluated with context spanning multiple utterances, have been shown to consistently outperform both conventional n-gram language mod…
Transformer-Transducers for Code-Switched Speech Recognition
Siddharth Dalmia, Yuzong Liu, Srikanth Ronanki +1
We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation…
Robust Prediction of Punctuation and Truecasing for Medical ASR
Monica Sunkara, Srikanth Ronanki, Kalpit Dixit +2
Automatic speech recognition (ASR) systems in the medical domain that focus on transcribing clinical dictations and doctor-patient conversations often pose many challenges due to t…
In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data
Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5
Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…