9 citations · 18 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2020
Answer-checking in Context: A Multi-modal FullyAttention Network for Visual Question Answering
Hantao Huang, Tao Han, Wei Han +2
Visual Question Answering (VQA) is challenging due to the complex cross-modal relations. It has received extensive attention from the research community. From the human perspective…
cs.CV2020★ 9 cited
Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering
Wei Han, Hantao Huang, Tao Han
Image text carries essential information to understand the scene and perform reasoning. Text-based visual question answering (text VQA) task focuses on visual questions that requir…