4 papers
GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs
Junqing Du, Fernando Ropero, Erkin Turkoz +2
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at…
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
Fernando Ropero, Erkin Turkoz, Daniel Matos +6
Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approac…
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
Antonio Ruiz, Tao Wu, Andrew Melnik +6
Methods that synthesize indoor 3D scenes from text prompts have wide-ranging applications in film production, interior design, video games, virtual reality, and synthetic data gene…
Lane Graph Extraction from Aerial Imagery via Lane Segmentation Refinement with Diffusion Models
Antonio Ruiz, Andrew Melnik, Nicolo Savioli +3
The lane graph is critical for applications such as autonomous driving and lane-level route planning. While previous research has focused on extracting lane-level graphs from aeria…