1 paper
Gabriele Lombardo, Luigi Maiorana, Liliana Lo Presti +1
Visual Grounding benchmarks assume that the object described by a referring expression is always present in the image, and grounding models are therefore rarely evaluated under sem…