5 papers
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, boundary conditions, and overlap.…
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present a…
Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
We describe \emph{Dual-Route Top-K Retrieval with 1v1 VLM Reranking} for the CoVR-R challenge. The method treats composed video retrieval as two coupled problems: finding a suffici…
VEU-Bench: Towards Comprehensive Understanding of Video Editing
Bozheng Li, Yongliang Wu, Yi Lu +7
Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, thei…
Number it: Temporal Grounding Videos like Flipping Manga
Yongliang Wu, Xinting Hu, Yuyang Sun +5
Video Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding…