2 citations · 2 across the 2 of their papers we have counts for
3 papers
Representation learning through cross-modal conditional teacher-student training for speech emotion recognition
Sundararajan Srinivasan, Zhaocheng Huang, Katrin Kirchhoff
Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to eff…
Speaker-conversation factorial designs for diarization error analysis
Scott Seyfarth, Sundararajan Srinivasan, Katrin Kirchhoff
Speaker diarization accuracy can be affected by both acoustics and conversation characteristics. Determining the cause of diarization errors is difficult because speaker voice acou…
Best of Both Worlds: Robust Accented Speech Recognition with Adversarial Transfer Learning
Nilaksh Das, Sravan Bodapati, Monica Sunkara +2
Training deep neural networks for automatic speech recognition (ASR) requires large amounts of transcribed speech. This becomes a bottleneck for training robust models for accented…