4 papers · 1 filter
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
Lars Doorenbos, Duc Manh Vu, Serdar Ozsoy +1
The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this additional information can…
The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2
Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…
Video Panels for Long Video Understanding
Lars Doorenbos, Federico Spurio, Juergen Gall
Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short…
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos +3
Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work…