3 papers
cs.CV2026
Towards Long-Form Spatio-Temporal Video Grounding
Xin Gu, Bing Fan, Jiali Yao +5
In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on loc…
cs.CV2025
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
Libo Zhang, Yongsheng Yu, Jiali Yao +1
Generative Adversarial Network (GAN) inversion have demonstrated excellent performance in image inpainting that aims to restore lost or damaged image texture using its unmasked con…
cs.CV2025
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
Jiali Yao, Xinran Deng, Xin Gu +6
In this paper, we propose spatio-temporal omni-object video grounding, dubbed OmniSTVG, a new STVG task that aims at localizing spatially and temporally all targets mentioned in th…