5 papers
On Test-Time Scaling for Vision-Language Models
Fawaz Sammani, Tzoulio Chamiti, Nikos Deligiannis
Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. While it has been widely studi…
When Negation Is a Geometry Problem in Vision-Language Models
Fawaz Sammani, Tzoulio Chamiti, Paul Gavrikov +1
Joint Vision-Language Embedding models such as CLIP typically fail at understanding negation in text queries, for example, failing to distinguish "no" in the query: "a plain blue s…
CLIP-Free, Label Free, Unsupervised Concept Bottleneck Models
Fawaz Sammani, Jonas Fischer, Nikos Deligiannis
Concept Bottleneck Models (CBMs) map dense feature representations into human-interpretable concepts which are then combined linearly to make a prediction. However, modern CBMs rel…
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
Ada Gorgun, Fawaz Sammani, Nikos Deligiannis +2
Diffusion models are usually evaluated by their final outputs, gradually denoising random noise into meaningful images. Yet, generation unfolds along a trajectory, and analyzing th…
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
Fawaz Sammani, Nikos Deligiannis
Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retriev…