1 paper
Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang +2
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large l…