activity
20242026
most citedSuFIA: Language-Guided Augmented Dexterity for Robotic Surgical Assistants

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Post-Training VLMs for Video Mistake Detection

Federico Spurio, Olga Zatsarynna, Lars Doorenbos +3

Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecti…

cs.CV2026

Multi-Person Human Motion Forecasting in Complex Scenes

Serdar Ozsoy, Lars Doorenbos, Juergen Gall

Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorpora…

cs.CV2026

Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition

Lars Doorenbos, Duc Manh Vu, Serdar Ozsoy +1

The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this additional information can…

cs.CV2026

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2

Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…

cs.CV2025

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos +3

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work…

cs.CV2025

Video Panels for Long Video Understanding

Lars Doorenbos, Federico Spurio, Juergen Gall

Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short…