7 papers
Towards Long-Form Spatio-Temporal Video Grounding
Xin Gu, Bing Fan, Jiali Yao +5
In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on loc…
AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation
Haobo Li, Yanhong Zeng, Yunhong Lu +6
We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt adversarial distillation bu…
Towards Visual Query Localization in the 3D World
Liang Peng, Bohan Tan, Zhipeng Zhang +4
Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual q…
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
Xin Gu, Haoji Zhang, Qihang Fan +7
Spatio-temporal video grounding (STVG) requires localizing a target object in untrimmed videos both temporally and spatially from natural language descriptions. Despite their stron…
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
Jiali Yao, Xinran Deng, Xin Gu +6
In this paper, we propose spatio-temporal omni-object video grounding, dubbed OmniSTVG, a new STVG task that aims at localizing spatially and temporally all targets mentioned in th…
Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
Liting Lin, Heng Fan, Zhipeng Zhang +3
Motivated by the Parameter-Efficient Fine-Tuning (PEFT) in large language models, we propose LoRAT, a method that unveils the power of large ViT model for tracking within laborator…