4 papers
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
Sepehr Kazemi Ranjbar, Kumail Alhamoud, Marzyeh Ghassemi
Vision-Language Models (VLMs) struggle with negation. Given a prompt like "retrieve (or generate) a street scene without pedestrians," they often fail to respect the "not." Existin…
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
Ali Rasekh, Sepehr Kazemi Ranjbar, Simon Gottschalk
Explainable object recognition using vision-language models such as CLIP involves predicting accurate category labels supported by rationales that justify the decision-making proce…
ExIQA: Explainable Image Quality Assessment Using Distortion Attributes
Sepehr Kazemi Ranjbar, Emad Fatemizadeh
Blind Image Quality Assessment (BIQA) aims to develop methods that estimate the quality scores of images in the absence of a reference image. In this paper, we approach BIQA from a…
ECOR: Explainable CLIP for Object Recognition
Ali Rasekh, Sepehr Kazemi Ranjbar, Milad Heidari +1
Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. Their open vo…