4 papers · 1 filter
Cross-Attentive Multiview Fusion of Vision-Language Embeddings
Tomas Berriel Martins, Martin R. Oswald, Javier Civera
Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challengin…
Open-Vocabulary Online Semantic Mapping for SLAM
Tomas Berriel Martins, Martin R. Oswald, Javier Civera
This paper presents an Open-Vocabulary Online 3D semantic mapping pipeline, that we denote by its acronym OVO. Given a sequence of posed RGB-D frames, we detect and track 3D segmen…
Feature Splatting for Better Novel View Synthesis with Low Overlap
T. Berriel Martins, Javier Civera
3D Gaussian Splatting has emerged as a very promising scene representation, achieving state-of-the-art quality in novel view synthesis significantly faster than competing alternati…
Ray-Patch: An Efficient Querying for Light Field Transformers
T. Berriel Martins, Javier Civera
In this paper we propose the Ray-Patch querying, a novel model to efficiently query transformers to decode implicit representations into target views. Our Ray-Patch decoding reduce…