3 citations · 3 across the 4 of their papers we have counts for
1 paper · 1 filter
Ting Liu, Xuyang Liu, Siteng Huang +5
Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved perf…