4 papers
Towards Unconstrained Human-Object Interaction
Francesco Tonini, Alessandro Conti, Lorenzo Vaquero +2
Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on…
CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection
Andrea Appiani, Cigdem Beyan
Voice Activity Detection (VAD) is the process of automatically determining whether a person is speaking and identifying the timing of their speech in an audiovisual data. Tradition…
AL-GTD: Deep Active Learning for Gaze Target Detection
Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero +2
Gaze target detection aims at determining the image location where a person is looking. While existing studies have made significant progress in this area by regressing accurate ga…
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
Anil Osman Tur, Alessandro Conti, Cigdem Beyan +5
In smart retail applications, the large number of products and their frequent turnover necessitate reliable zero-shot object classification methods. The zero-shot assumption is ess…