2 papers
cs.CV2026
GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs
Junqing Du, Fernando Ropero, Erkin Turkoz +2
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at…
cs.CV2026
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
Fernando Ropero, Erkin Turkoz, Daniel Matos +6
Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approac…