1 citations · 1 across the 1 of their papers we have counts for
1 paper
Zhihong Chen, Ruifei Zhang, Yibing Song +2
Visual grounding (VG) aims to establish fine-grained alignment between vision and language. Ideally, it can be a testbed for vision-and-language models to evaluate their understand…