activity
20242026
collaborators

8 papers

cs.CV2026

Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

Niels Sombekke, Rob G. J. Wijnhoven, Martin R. Oswald

We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared…

cs.CV2026

Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery

Luuk Versteeg, Rob G. J. Wijnhoven, Martin R. Oswald

We present a method for jointly predicting instance-level roof segment masks together with three continuous geometric attributes -- building height, roof slope, and roof azimuth --…

cs.CV2026

Cross-Attentive Multiview Fusion of Vision-Language Embeddings

Tomas Berriel Martins, Martin R. Oswald, Javier Civera

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challengin…

cs.CV2025

Edge-Centric Relational Reasoning for 3D Scene Graph Prediction

Yanni Ma, Hao Liu, Yulan Guo +2

3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adop…

cs.CV2025

3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation

Weijie Wei, Osman Ülger, Fatemeh Karimi Nejadasl +2

Open-Vocabulary Segmentation (OVS) methods offer promising capabilities in detecting unseen object categories, but the category must be known and needs to be provided by a human, e…

cs.CV2025

SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation

Duy-Kien Nguyen, Martin R. Oswald, Cees G. M. Snoek

The ability to detect objects in images at varying scales has played a pivotal role in the design of modern object detectors. Despite considerable progress in removing hand-crafted…