18 citations · 28 across the 5 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2020★ 18 cited
Contrastive Visual-Linguistic Pretraining
Lei Shi, Kai Shuang, Shijie Geng +6
Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-l…
cs.CV2020★ 3 cited
Multi-Layer Content Interaction Through Quaternion Product For Visual Question Answering
Lei Shi, Shijie Geng, Kai Shuang +4
Multi-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-a…