6 papers · 1 filter
Multi-Person Human Motion Forecasting in Complex Scenes
Serdar Ozsoy, Lars Doorenbos, Juergen Gall
Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorpora…
Learning Probabilistic Embeddings for Unsupervised Action Segmentation
Shuai Li, Duc Manh Vu, Juergen Gall
This paper concerns the problem of unsupervised temporal action segmentation for long, untrimmed videos. Recent successful approaches follow a joint representation learning and clu…
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
Lars Doorenbos, Duc Manh Vu, Serdar Ozsoy +1
The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this additional information can…
The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
Serdar Ozsoy, Lars Doorenbos, Federico Spurio +2
Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the rea…
Video Panels for Long Video Understanding
Lars Doorenbos, Federico Spurio, Juergen Gall
Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short…
Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
Uzay Gökay, Federico Spurio, Dominik R. Bach +1
Current state-of-the-art methods for skeleton-based temporal action segmentation are predominantly supervised and require annotated data, which is expensive to collect. In contrast…