2 papers
cs.CL2023
End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu +5
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversat…
eess.AS2022
Representation learning through cross-modal conditional teacher-student training for speech emotion recognition
Sundararajan Srinivasan, Zhaocheng Huang, Katrin Kirchhoff
Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to eff…