18 citations · 21 across the 2 of their papers we have counts for
3 papers
cs.CV2020★ 3 cited
Overcoming Language Priors with Self-supervised Learning for Visual Question Answering
Xi Zhu, Zhendong Mao, Chunxiao Liu +3
Most Visual Question Answering (VQA) models suffer from the language prior problem, which is caused by inherent data biases. Specifically, VQA models tend to answer questions (e.g.…
cs.CV2020
Graph Structured Network for Image-Text Matching
Chunxiao Liu, Zhendong Mao, Tianzhu Zhang +3
Image-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works…
cs.MM2019★ 18 cited
Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching
Chunxiao Liu, Zhendong Mao, An-An Liu +3
Learning semantic correspondence between image and text is significant as it bridges the semantic gap between vision and language. The key challenge is to accurately find and corre…