1 paper
Ling Li, Bowen Liu, Zinuo Zhan +4
Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores…