2 citations · 2 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2019
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Gen Li, Nan Duan, Yuejian Fang +3
We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained mo…
cs.CV2019★ 2 cited
Deep Reason: A Strong Baseline for Real-World Visual Reasoning
Chenfei Wu, Yanzhao Zhou, Gen Li +3
This paper presents a strong baseline for real-world visual reasoning (GQA), which achieves 60.93% in GQA 2019 challenge and won the sixth place. GQA is a large dataset with 22M qu…