3 papers
cs.CV2026
Towards Spatial Supersensing in the Wild
Tianjun Gu, Tianyu Xin, Kuan Zhang +12
The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…
cs.AI2026
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners
Zheng Lu, Mingqi Gao, Qinlei Xie +8
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic…
cs.CV2026
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
Kuan Zhang, Dongchen Liu, Qiyue Zhao +5
Human gameplay is a visually grounded interaction loop in which players act, reflect on failures, and watch tutorials to refine strategies. Can Vision-Language Models (VLMs) also l…