1 citations · 1 across the 10 of their papers we have counts for
1 paper · 1 filter
Hongbing Li, Linhui Xiao, Zihan Zhao +4
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent…