8 papers
SoccerNet 2026 Challenges Results
Anthony Cioppa, Silvio Giancola, Håkan Ardö +102
The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video underst…
Inferring Dynamic Physical Properties from Video Foundation Models
Guanqi Zhan, Xianzheng Ma, Weidi Xie +1
We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasti…
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
Guanqi Zhan, Yuanpei Liu, Kai Han +2
The objective in this paper is to improve the performance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-tr…
Character-Centric Understanding of Animated Movies
Zhongrui Gui, Junyu Xie, Tengda Han +2
Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the c…
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
Junyu Xie, Tengda Han, Max Bain +5
Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework tha…
Aerial Monocular 3D Object Detection
Yue Hu, Shaoheng Fang, Weidi Xie +1
Drones equipped with cameras can significantly enhance human ability to perceive the world because of their remarkable maneuverability in 3D space. Ironically, object detection for…