19 citations · 19 across the 6 of their papers we have counts for
6 papers
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
Navid Rajabi, Jana Kosecka
Vision and Language Models (VLMs) continue to demonstrate remarkable zero-shot (ZS) performance across various tasks. However, many probing studies have revealed that even the best…
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Ivana Beňová, Jana Košecká, Michal Gregor +3
The dominant probing approaches rely on the zero-shot performance of image-text matching tasks to gain a finer-grained understanding of the representations learned by recent multim…
Labeling Indoor Scenes with Fusion of Out-of-the-Box Perception Models
Yimeng Li, Navid Rajabi, Sulabh Shrestha +2
The image annotation stage is a critical and often the most time-consuming part required for training and evaluating object detection and semantic segmentation models. Deployment o…
Graph-CoVis: GNN-based Multi-view Panorama Global Pose Estimation
Negar Nejatishahidin, Will Hutchcroft, Manjunath Narayana +5
In this paper, we address the problem of wide-baseline camera pose estimation from a group of 360 panoramas under upright-camera assumption. Recent work has demonstrated th…
U2RLE: Uncertainty-Guided 2-Stage Room Layout Estimation
Pooya Fayyazsanavi, Zhiqiang Wan, Will Hutchcroft +4
While the existing deep learning-based room layout estimation techniques demonstrate good overall accuracy, they are less effective for distant floor-wall boundary. To tackle this…
Semantic Image Based Geolocation Given a Map
Arsalan Mousavian, Jana Kosecka
The problem visual place recognition is commonly used strategy for localization. Most successful appearance based methods typically rely on a large database of views endowed with l…