4 papers
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
Zhaoyu Liu, Xi Weng, Lianyu Hu +4
Tennis is one of the most widely followed sports, generating extensive broadcast footage with strong potential for professional analysis, automated coaching, and real-time commenta…
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Yuxuan Wang, Yiqi Song, Cihang Xie +2
Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demand…
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
Zhitao Zeng, Zhu Zhuo, Xiaojun Jia +12
Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remain…