activity
20122022
most citedSelf-Attention Networks for Connectionist Temporal Classification in Speech Recognition

133 citations · 226 across the 12 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2022

Representation learning through cross-modal conditional teacher-student training for speech emotion recognition

Sundararajan Srinivasan, Zhaocheng Huang, Katrin Kirchhoff

Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to eff…

eess.AS2021

Remember the context! ASR slot error correction through memorization

Dhanush Bekal, Ashish Shenoy, Monica Sunkara +2

Accurate recognition of slot values such as domain specific words or named entities by automatic speech recognition (ASR) systems forms the core of the Goal-oriented Dialogue Syste…

eess.AS20214 cited

ASR Adaptation for E-commerce Chatbots using Cross-Utterance Context and Multi-Task Language Modeling

Ashish Shenoy, Sravan Bodapati, Katrin Kirchhoff

Automatic Speech Recognition (ASR) robustness toward slot entities are critical in e-commerce voice assistants that involve monetary transactions and purchases. Along with effectiv…

eess.AS2021

Speaker-conversation factorial designs for diarization error analysis

Scott Seyfarth, Sundararajan Srinivasan, Katrin Kirchhoff

Speaker diarization accuracy can be affected by both acoustics and conversation characteristics. Determining the cause of diarization errors is difficult because speaker voice acou…

eess.AS2020

Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment

Ethan A. Chi, Julian Salazar, Katrin Kirchhoff

Non-autoregressive models greatly improve decoding speed over typical sequence-to-sequence models, but suffer from degraded performance. Infilling and iterative refinement models m…

eess.AS2020

Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech

Monica Sunkara, Srikanth Ronanki, Dhanush Bekal +2

In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data.…