33 citations · 33 across the 1 of their papers we have counts for
1 paper
Zhiling Yan, Kai Zhang, Rong Zhou +3
In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA)…