5 papers
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
Jian Hu, Zixu Cheng, Shaogang Gong +4
Video Temporal Grounding (TG) aims to temporally locate video segments matching a natural language description (a query) in a long video. While Vision-Language Models (VLMs) are ef…
ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
Jian Hu, Dimitrios Korkinof, Shaogang Gong +1
We introduce ViSMap: Unsupervised Video Summarisation by Meta Prompting, a system to summarise hour long videos with no-supervision. Most existing video understanding models work w…
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Zixu Cheng, Jian Hu, Ziquan Liu +3
Human processes video reasoning in a sequential spatio-temporal reasoning logic, we first identify the relevant frames ("when") and then analyse the spatial relationships ("where")…
CoS: Chain-of-Shot Prompting for Long Video Understanding
Jian Hu, Zixu Cheng, Chenyang Si +2
Multi-modal Large Language Models (MLLMs) struggle with long videos due to the need for excessive visual tokens. These tokens exceed massively the context length of MLLMs, resultin…
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
Jian Hu, Zixu Cheng, Shaogang Gong
Task-generic promptable image segmentation aims to achieve segmentation of diverse samples under a single task description by utilizing only one task-generic prompt. Current method…