1 paper
Zhecan Wang, Junzhang Liu, Chia-Wei Tang +11
Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perfor…