collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

Post-Training VLMs for Video Mistake Detection

Federico Spurio, Olga Zatsarynna, Lars Doorenbos +3

Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecti…

cs.CV2026

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

Di Yang, Mahmoud Ali, Quan Kong +2

Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing models rely primarily on appeara…

cs.CV2026

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2

Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…

cs.CV2026

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Alessandro Simoni, Riccardo Catalini, Davide Di Nucci +6

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular,…

cs.CV2025

Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization

Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta +3

In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuo…

cs.CV2025

Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions

Federico Spurio, Emad Bahrami, Olga Zatsarynna +3

We introduce Action Discovery, a novel setup within Temporal Action Segmentation that addresses the challenge of defining and annotating ambiguous actions and incomplete annotation…