5 papers
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
Francesco Gentile, Nicola Dall'Asen, Francesco Tonini +3
As vision-language models are deployed at scale, understanding their internal mechanisms becomes increasingly critical. Existing interpretability methods predominantly rely on acti…
MAMBO: High-Resolution Generative Approach for Mammography Images
Milica Å kipina, Nikola JoviÅ¡iÄ, Nicola Dall'Asen +5
Mammography is the gold standard for the detection and diagnosis of breast cancer. This procedure can be significantly enhanced with Artificial Intelligence (AI)-based software, wh…
Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
Francesco Tonini, Lorenzo Vaquero, Alessandro Conti +2
Human-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets wit…
AL-GTD: Deep Active Learning for Gaze Target Detection
Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero +2
Gaze target detection aims at determining the image location where a person is looking. While existing studies have made significant progress in this area by regressing accurate ga…
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
Anil Osman Tur, Alessandro Conti, Cigdem Beyan +5
In smart retail applications, the large number of products and their frequent turnover necessitate reliable zero-shot object classification methods. The zero-shot assumption is ess…