2 papers
cs.CV2024
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
Navid Rajabi, Jana Kosecka
Vision and Language Models (VLMs) continue to demonstrate remarkable zero-shot (ZS) performance across various tasks. However, many probing studies have revealed that even the best…
cs.CV2023
Labeling Indoor Scenes with Fusion of Out-of-the-Box Perception Models
Yimeng Li, Navid Rajabi, Sulabh Shrestha +2
The image annotation stage is a critical and often the most time-consuming part required for training and evaluating object detection and semantic segmentation models. Deployment o…