3 papers
cs.RO2026
-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Xiaowei Cai, Yunuo Cai, Bingao Chen +36
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-acti…
cs.CV2025
AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
Junwen Miao, Penghui Du, Yingying Fan +5
Industrial anomaly detection (IAD) is challenging due to the subtle and highly localized nature of many defects, which single-pass vision--language models (VLMs) often fail to capt…
cs.CV2025
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
Haotian Xia, Haonan Ge, Junbo Zou +16
Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models…