#spatial reasoning
13 resultsSpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications
Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1
The paper introduces ByDeWay-V2, a training‑free prompting framework that adds explicit pairwise spatial predicates derived from depth estimation and open‑vocabulary object detecti…
Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci
The paper presents Rad-JEPA 3D, a self‑supervised joint‑embedding model that learns 3D CT representations by predicting latent features of a full scan from a masked view, using a h…
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
Quoc-Huy Trinh, Xi Ding, Yang Liu +7
The paper introduces SpatialMed, a benchmark and an automated pipeline that generates 3D spatial visual question‑answer pairs for medical imaging, and shows that current multimodal…
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
The paper introduces ViPS, a framework that combines multiple visual priors from diverse foundation models using an Efficient Prior Proxy and Dynamic Prior Fusion to improve spatia…
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6
The paper proposes Ego Scene Augmentation (ESA), a framework that uses an Ego-element Graph to improve the spatial perception of multimodal large language models for egocentric vis…
Towards Spatial Supersensing in the Wild
Tianjun Gu, Tianyu Xin, Kuan Zhang +12
The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5
The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…
S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving
Jianguo Yu, Rukang Wang, Duanfeng Chu +3
The paper introduces S-squared-VLA, a vision‑language‑action model that separates semantic intent reasoning from spatial geometry processing to improve low‑level control for autono…
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
Xinyue Xu, Zheng Zhang, Kunyang Ma +5
The paper introduces DM-KG, a direction‑metric knowledge graph that extracts 3D spatial relationships from street‑view images and injects them into vision‑language models to improv…
Egocentric Bias in Vision-Language Models
Maijunxian Wang, Yijiang Li, Bingyang Wang +6
The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…
LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning
Shiyi Mu, Xujie Chen, Shugong Xu
The paper introduces LARAD, a method for detecting road anomalies in autonomous driving by training models to recognize spatial‑logic violations rather than relying on texture diff…
When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning
Quynh Vo, Phuc Dao, Cong-Duy Nguyen +1
The paper introduces Depth-Ordinal Prompting (DOP), a training‑free technique that converts monocular depth estimates into object‑level ordinal text cues, enabling vision‑language…