2 citations · 3 across the 14 of their papers we have counts for
10 papers · 1 filter
MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
Luca Collorone, Matteo Gioia, Massimiliano Pappa +5
Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, e…
Human Motion Unlearning
Edoardo De Matteis, Matteo Migliarini, Alessio Sampieri +2
We introduce Human Motion Unlearning and motivate it through the concrete task of preventing violent 3D motion synthesis, an important safety requirement given that popular text-to…
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
Luca Collorone, Stefano D'Arrigo, Massimiliano Pappa +3
We introduce the novel task of Crowd Volume Estimation (CVE), defined as the process of estimating the collective body volume of crowds using only RGB images. Besides event managem…
Social EgoMesh Estimation
Luca Scofano, Alessio Sampieri, Edoardo De Matteis +2
Accurately estimating the 3D pose of the camera wearer in egocentric video sequences is crucial to modeling human behavior in virtual and augmented reality applications. The task p…
TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
Leonardo Plini, Luca Scofano, Edoardo De Matteis +6
Identifying procedural errors online from egocentric videos is a critical yet challenging task across various domains, including manufacturing, healthcare, and skill-based training…
Compositional Entailment Learning for Hyperbolic Vision-Language Models
Avik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno +3
Image-text representation learning forms a cornerstone in vision-language models, where pairs of images and textual descriptions are contrastively aligned in a shared embedding spa…