6 papers
IAE-VTG: Interaction-Aligned Action-Entity Video Temporal Grounding
Shiwen Zhao, Qi Zhang, Sezer Karaoglu +2
Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a particular entity. Existing methods…
Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery
Niels Sombekke, Rob G. J. Wijnhoven, Martin R. Oswald
We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared…
Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery
Luuk Versteeg, Rob G. J. Wijnhoven, Martin R. Oswald
We present a method for jointly predicting instance-level roof segment masks together with three continuous geometric attributes -- building height, roof slope, and roof azimuth --…
Cross-Attentive Multiview Fusion of Vision-Language Embeddings
Tomas Berriel Martins, Martin R. Oswald, Javier Civera
Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challengin…
Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
Yanni Ma, Hao Liu, Yulan Guo +2
3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adop…
3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation
Weijie Wei, Osman Ülger, Fatemeh Karimi Nejadasl +2
Open-Vocabulary Segmentation (OVS) methods offer promising capabilities in detecting unseen object categories, but the category must be known and needs to be provided by a human, e…