11 papers
Towards Comprehensive Basketball Understanding
Yirong Hu, Jiayuan Rao, Yu Zhang +2
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily…
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Zhaokai Wang, Tianlin Gui, Jiayuan Rao +3
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available…
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
Yibin Yan, Jilan Xu, Shangzhe Di +2
Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation…
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Yikun Liu, Yuan Liu, Shangzhe Di +8
Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
Yudi Shi, Shangzhe Di, Qirui Chen +5
Video reasoning constitutes a comprehensive assessment of a model's capabilities, as it demands robust perceptual and interpretive skills, thereby serving as a means to explore the…
Revisiting Multi-Task Visual Representation Learning
Shangzhe Di, Zhonghua Zhai, Weidi Xie
Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised…