2 papers
cs.CL2022
Cross-stitched Multi-modal Encoders
Karan Singla, Daniel Pressel, Ryan Price +3
In this paper, we propose a novel architecture for multi-modal speech and text input. We combine pretrained speech and text encoders using multi-headed cross-modal attention and jo…
cs.CL2022
Seq-2-Seq based Refinement of ASR Output for Spoken Name Capture
Karan Singla, Shahab Jalalvand, Yeon-Jun Kim +3
Person name capture from human speech is a difficult task in human-machine conversations. In this paper, we propose a novel approach to capture the person names from the caller utt…