collaborators

6 papers

cs.CV2026

SoccerNet 2026 Challenges Results

Anthony Cioppa, Silvio Giancola, Håkan Ardö +102

The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video underst…

cs.CV2026

OneFocus: Enabling Real-World X-ray Security Screening with a Unified Vision-Language Model

Jiali Wen, Hongxia Gao, Litao Li +4

X-ray contraband detection is critical for security in large-scale logistics and transportation, yet conventional detectors struggle to adapt to emerging contraband types and lack…

cs.CV2026

MSUE: Multi-Modal Soccer Understanding Expert

Litao Li, Yibo Yu, Yufeng Hu +4

This paper presents our solution to the 2026 SoccerNet VQA Challenge. We first develop a cost-effective data synthesis pipeline driven by a Vision-Language Model (VLM), which syste…

cs.CV2026

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

Junchao Liao, Zhenghao Zhang, Xiangyu Meng +5

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains c…

cs.CV2026

XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening

Hongxia Gao, Litao Li, Yixin Chen +3

X-ray contraband detection is critical for public safety. However, current methods primarily rely on bounding box annotations, which limit model generalization and performance due…

cs.AI2026

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

Lakshya Gupta, Litao Li, Yizhe Liu +5

Frontier large language models (LLMs) excel as autonomous agents in many domains, yet they remain untested in complex enterprise systems where hidden workflows create cascading eff…