4 citations · 4 across the 2 of their papers we have counts for
7 papers
When Negation Is a Geometry Problem in Vision-Language Models
Fawaz Sammani, Tzoulio Chamiti, Paul Gavrikov +1
Joint Vision-Language Embedding models such as CLIP typically fail at understanding negation in text queries, for example, failing to distinguish "no" in the query: "a plain blue s…
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
Ada Gorgun, Fawaz Sammani, Nikos Deligiannis +2
Diffusion models are usually evaluated by their final outputs, gradually denoising random noise into meaningful images. Yet, generation unfolds along a trajectory, and analyzing th…
CLIP-Free, Label Free, Unsupervised Concept Bottleneck Models
Fawaz Sammani, Jonas Fischer, Nikos Deligiannis
Concept Bottleneck Models (CBMs) map dense feature representations into human-interpretable concepts which are then combined linearly to make a prediction. However, modern CBMs rel…
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
Fawaz Sammani, Nikos Deligiannis
Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retriev…
NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks
Fawaz Sammani, Tanmoy Mukherjee, Nikos Deligiannis
Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, hi…
Show, Edit and Tell: A Framework for Editing Image Captions
Fawaz Sammani, Luke Melas-Kyriazi
Most image captioning frameworks generate captions directly from images, learning a mapping from visual features to natural language. However, editing existing captions can be easi…