From the 1 of 4 linked papers with an AI index.
4 papers
Towards Spatial Supersensing in the Wild
Tianjun Gu, Tianyu Xin, Kuan Zhang +12
The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
Xiaolong Zhou, Yifei Liu, Ziyang Gong +8
Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pristine visual inputs and overl…
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
Kuan Zhang, Dongchen Liu, Qiyue Zhao +12
The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this singular physical existence…
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
Kuan Zhang, Dongchen Liu, Qiyue Zhao +5
Human gameplay is a visually grounded interaction loop in which players act, reflect on failures, and watch tutorials to refine strategies. Can Vision-Language Models (VLMs) also l…