activity
20242026
most citedDynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

2 citations · 2 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CV2026

FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani +3

Human-Object Interaction (HOI) detection aims to localize humans and objects in images and classify their interactions. Zero-shot HOI focuses on recognizing interactions that are n…

cs.CV2026

Zero-Shot Temporal Action Localization Through Textual Guidance

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero +3

Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work…

cs.CV2026

Training-Free Semantic Multi-Object Tracking with Vision-Language Models

Laurence Bonat, Francesco Tonini, Elisa Ricci +1

Semantic Multi-Object Tracking (SMOT) extends multi-object tracking with semantic outputs such as video summaries, instance-level captions, and interaction labels, aiming to move f…

cs.CV2026

Towards Unconstrained Human-Object Interaction

Francesco Tonini, Alessandro Conti, Lorenzo Vaquero +2

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on…

cs.CV2026

From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition

Francesco Gentile, Nicola Dall'Asen, Francesco Tonini +3

As vision-language models are deployed at scale, understanding their internal mechanisms becomes increasingly critical. Existing interpretability methods predominantly rely on acti…

cs.CV2025

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero +3

What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where t…