1 paper
Xingjian Wang, Shijian Wang, Yibo Wang +4
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely ov…