1 paper · 1 filter
Leqian Ding, Junning Qiu, Manwen Yang +2
Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progres…