1 paper
Haoxuan Chen, Xianqin Liu, Jian-Fang Hu
Spatio-Temporal Video Grounding aims to localize object tubes based on textual queries. While recent methods have achieved remarkable success, they mainly focus on high-quality(HQ)…