2 citations · 3 across the 2 of their papers we have counts for
3 papers
eess.AS2022★ 1 cited
Monolingual Recognizers Fusion for Code-switching Speech Recognition
Tongtong Song, Qiang Xu, Haoyu Lu +5
The bi-encoder structure has been intensively investigated in code-switching (CS) automatic speech recognition (ASR). However, most existing methods require the structures of two m…
cs.CV2022★ 2 cited
COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval
Haoyu Lu, Nanyi Fei, Yuqi Huo +3
Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recentl…
cs.CV2021
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…