22 citations · 30 across the 5 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 6 cited
HiVLP: Hierarchical Vision-Language Pre-Training for Fast Image-Text Retrieval
Feilong Chen, Xiuyi Chen, Jiaxin Shi +3
In the past few years, the emergence of vision-language pre-training (VLP) has brought cross-modal retrieval to a new era. However, due to the latency and computation demand, it is…
cs.CV2022
Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning
Feilong Chen, Xiuyi Chen, Shuang Xu +1
Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog histo…