27 citations · 51 across the 30 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023
Generalized zero-shot audio-to-intent classification
Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi +2
Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generali…
cs.SD2023
Masked Audio Text Encoders are Effective Multi-Modal Rescorers
Jinglun Cai, Monica Sunkara, Xilai Li +3
Masked Language Models (MLMs) have proven to be effective for second-pass rescoring in Automatic Speech Recognition (ASR) systems. In this work, we propose Masked Audio Text Encode…