8 papers
SoccerNet 2026 Challenges Results
Anthony Cioppa, Silvio Giancola, Håkan Ardö +102
The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video underst…
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, boundary conditions, and overlap.…
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present a…
Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
We describe \emph{Dual-Route Top-K Retrieval with 1v1 VLM Reranking} for the CoVR-R challenge. The method treats composed video retrieval as two coupled problems: finding a suffici…
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
Yuechen Luo, Fang Li, Shaoqing Xu +10
While Vision-Language-Action (VLA) models have revolutionized autonomous driving by unifying perception and planning, their reliance on explicit textual Chain-of-Thought (CoT) lead…
OpusAnimation: Code-Based Dynamic Chart Generation
Bozheng Li, Miao Yang, Zhenhan Chen +9
Dynamic Chart Generation (DCG) involves producing code-rendered animated visualizations as charts. While recent advances in multi-modal large language models (MLLMs) have significa…