#spatial reasoning

13 results
cs.AI2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Yang Zhou, Zixuan Huang, Sunzhu Li +10

The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…

#vision-language models#spatial reasoning#tool use#embodied AI
cs.CV2026

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1

The paper introduces ByDeWay-V2, a training‑free prompting framework that adds explicit pairwise spatial predicates derived from depth estimation and open‑vocabulary object detecti…

#multimodal large language models#spatial reasoning#depth estimation#prompt engineering
cs.CV2026

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci

The paper presents Rad-JEPA 3D, a self‑supervised joint‑embedding model that learns 3D CT representations by predicting latent features of a full scan from a masked view, using a h…

#self-supervised learning#3d medical imaging#ct scan analysis#joint embedding
cs.CV2026

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

Quoc-Huy Trinh, Xi Ding, Yang Liu +7

The paper introduces SpatialMed, a benchmark and an automated pipeline that generates 3D spatial visual question‑answer pairs for medical imaging, and shows that current multimodal…

#medical imaging#multimodal large language models#spatial reasoning#3d vision
cs.CV2026

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

Xiao Lin, Xiaohu Huang, Kai Han

The paper introduces ViPS, a framework that combines multiple visual priors from diverse foundation models using an Efficient Prior Proxy and Dynamic Prior Fusion to improve spatia…

#multimodal large language models#visual priors#spatial reasoning#prior fusion
cs.CV2026

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6

The paper proposes Ego Scene Augmentation (ESA), a framework that uses an Ego-element Graph to improve the spatial perception of multimodal large language models for egocentric vis…

#egocentric vision#visual question answering#spatial reasoning#multimodal language models
cs.CV2026

Towards Spatial Supersensing in the Wild

Tianjun Gu, Tianyu Xin, Kuan Zhang +12

The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…

#spatial reasoning#long-term video understanding#multimodal world modeling#benchmark dataset
cs.CV2026

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5

The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…

#3d medical imaging#multimodal large language models#spatial reasoning#chain-of-thought reasoning
cs.RO2026

S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

Jianguo Yu, Rukang Wang, Duanfeng Chu +3

The paper introduces S-squared-VLA, a vision‑language‑action model that separates semantic intent reasoning from spatial geometry processing to improve low‑level control for autono…

#autonomous driving#vision-language models#spatial reasoning#dual-stream architecture
cs.CV2026

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

Xinyue Xu, Zheng Zhang, Kunyang Ma +5

The paper introduces DM-KG, a direction‑metric knowledge graph that extracts 3D spatial relationships from street‑view images and injects them into vision‑language models to improv…

#spatial reasoning#vision-language models#street view imagery#knowledge graph
cs.CV2026

Egocentric Bias in Vision-Language Models

Maijunxian Wang, Yijiang Li, Bingyang Wang +6

The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…

#visual perspective taking#egocentric bias#vision-language models#spatial reasoning
cs.CV2026

LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning

Shiyi Mu, Xujie Chen, Shugong Xu

The paper introduces LARAD, a method for detecting road anomalies in autonomous driving by training models to recognize spatial‑logic violations rather than relying on texture diff…

#anomaly detection#autonomous driving#spatial reasoning#semantic segmentation
cs.CV2026

When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

Quynh Vo, Phuc Dao, Cong-Duy Nguyen +1

The paper introduces Depth-Ordinal Prompting (DOP), a training‑free technique that converts monocular depth estimates into object‑level ordinal text cues, enabling vision‑language…

#spatial reasoning#depth estimation#vision-language models#prompt engineering