133 citations · 226 across the 12 of their papers we have counts for
3 papers · 1 filter
Device Directedness with Contextual Cues for Spoken Dialog Systems
Dhanush Bekal, Sundararajan Srinivasan, Sravan Bodapati +2
In this work, we define barge-in verification as a supervised learning task where audio-only information is used to classify user spoken dialogue into true and false barge-ins. Fol…
Towards Personalization of CTC Speech Recognition Models with Contextual Adapters and Adaptive Boosting
Saket Dingliwal, Monica Sunkara, Sravan Bodapati +3
End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregr…
Representation learning through cross-modal conditional teacher-student training for speech emotion recognition
Sundararajan Srinivasan, Zhaocheng Huang, Katrin Kirchhoff
Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to eff…