3 papers
cs.CV2026
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
Vidi Team, Chia-Wen Kuo, Chuang Huang +31
Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to…
cs.CV2025
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
Xin Gu, Haoji Zhang, Qihang Fan +7
Spatio-temporal video grounding (STVG) requires localizing a target object in untrimmed videos both temporally and spatially from natural language descriptions. Despite their stron…
cs.CV2024
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
Yulin Wang, Haoji Zhang, Yang Yue +4
This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation com…