1 paper
Zihang Lin, Chaolei Tan, Jian-Fang Hu +3
In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which mode…