9 papers · 1 filter
Post-Training VLMs for Video Mistake Detection
Federico Spurio, Olga Zatsarynna, Lars Doorenbos +3
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecti…
T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition
Di Yang, Mahmoud Ali, Quan Kong +2
Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing models rely primarily on appeara…
The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2
Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…
SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
Alessandro Simoni, Riccardo Catalini, Davide Di Nucci +6
Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular,…
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
Sina Mokhtarzadeh Azar, Emad Bahrami, Enrico Pallotta +3
In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuo…
Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions
Federico Spurio, Emad Bahrami, Olga Zatsarynna +3
We introduce Action Discovery, a novel setup within Temporal Action Segmentation that addresses the challenge of defining and annotating ambiguous actions and incomplete annotation…