6 papers · 1 filter
SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport
Cigdem Beyan, Tonje Knutsen Sordalen, Kim Tallaksen Halvorsen
Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific body regions rather than distri…
Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification
Anil Osman Tur, Tonje Knutsen Sordalen, Kim Tallaksen Halvorsen +1
Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although recent vision-language models provide stron…
Towards Unconstrained Human-Object Interaction
Francesco Tonini, Alessandro Conti, Lorenzo Vaquero +2
Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on…
CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection
Andrea Appiani, Cigdem Beyan
Voice Activity Detection (VAD) is the process of automatically determining whether a person is speaking and identifying the timing of their speech in an audiovisual data. Tradition…
AL-GTD: Deep Active Learning for Gaze Target Detection
Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero +2
Gaze target detection aims at determining the image location where a person is looking. While existing studies have made significant progress in this area by regressing accurate ga…
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
Anil Osman Tur, Alessandro Conti, Cigdem Beyan +5
In smart retail applications, the large number of products and their frequent turnover necessitate reliable zero-shot object classification methods. The zero-shot assumption is ess…