5 citations · 10 across the 3 of their papers we have counts for
1 paper · 1 filter
Mengze Li, Tianbao Wang, Haoyu Zhang +6
Video Object Grounding (VOG) is the problem of associating spatial object regions in the video to a descriptive natural language query. This is a challenging vision-language task t…