1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Zaiquan Yang, Yuhao Liu, Gerhard Hancke +1
Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large lang…