activity
20182022
most citedIn Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2022

Device Directedness with Contextual Cues for Spoken Dialog Systems

Dhanush Bekal, Sundararajan Srinivasan, Sravan Bodapati +2

In this work, we define barge-in verification as a supervised learning task where audio-only information is used to classify user spoken dialogue into true and false barge-ins. Fol…

cs.CL2022

Towards Personalization of CTC Speech Recognition Models with Contextual Adapters and Adaptive Boosting

Saket Dingliwal, Monica Sunkara, Sravan Bodapati +3

End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregr…

cs.CL2021

Adapting Long Context NLM for ASR Rescoring in Conversational Agents

Ashish Shenoy, Sravan Bodapati, Monica Sunkara +2

Neural Language Models (NLM), when trained and evaluated with context spanning multiple utterances, have been shown to consistently outperform both conventional n-gram language mod…

cs.CL2020

Transformer-Transducers for Code-Switched Speech Recognition

Siddharth Dalmia, Yuzong Liu, Srikanth Ronanki +1

We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation…

cs.CL2020

Robust Prediction of Punctuation and Truecasing for Medical ASR

Monica Sunkara, Srikanth Ronanki, Kalpit Dixit +2

Automatic speech recognition (ASR) systems in the medical domain that focus on transcribing clinical dictations and doctor-patient conversations often pose many challenges due to t…

cs.CL20191 cited

In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5

Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…