collaborators

9 papers

cs.CV2026

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

Di Yang, Mahmoud Ali, Quan Kong +2

Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing models rely primarily on appeara…

cs.CV2026

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2

Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…

cs.CV2026

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Alessandro Simoni, Riccardo Catalini, Davide Di Nucci +6

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular,…

cs.CV2026

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation

Riccardo Catalini, Davide Di Nucci, Guido Borghi +5

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with un…

cs.CV2025

Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization

Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta +3

In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuo…

cs.CV2025

Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions

Federico Spurio, Emad Bahrami, Olga Zatsarynna +3

We introduce Action Discovery, a novel setup within Temporal Action Segmentation that addresses the challenge of defining and annotating ambiguous actions and incomplete annotation…