1 citations · 1 across the 1 of their papers we have counts for
1 paper
Siyu Zhang, Yeming Chen, Yaoru Sun +3
Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts h…