27 citations · 39 across the 4 of their papers we have counts for
1 paper · 1 filter
Jian Lu, Shikhar Srivastava, Junyu Chen +4
With the advent of multi-modal large language models (MLLMs), datasets used for visual question answering (VQA) and referring expression comprehension have seen a resurgence. Howev…