1 paper · 1 filter
Kai Chen, Ming Dai, Wenxuan Cheng +1
The paper introduces ScanFocus, a coarse-to-fine framework for spatio-temporal video grounding that first scans videos globally to generate coarse object proposals and then refines…