11 papers
Post-Training VLMs for Video Mistake Detection
Federico Spurio, Olga Zatsarynna, Lars Doorenbos +3
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecti…
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
Emad Bahrami, Olga Zatsarynna, Parth Pathak +3
We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for video question answering. While grou…
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
Syed Talal Wasim, Hamid Suleman, Olga Zatsarynna +2
We present MixANT, a novel architecture for stochastic long-term dense anticipation of human activities. While recent State Space Models (SSMs) like Mamba have shown promise throug…
Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions
Federico Spurio, Emad Bahrami, Olga Zatsarynna +3
We introduce Action Discovery, a novel setup within Temporal Action Segmentation that addresses the challenge of defining and annotating ambiguous actions and incomplete annotation…
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
Xuying Huang, Sicong Pan, Olga Zatsarynna +2
RGB-based semantic segmentation has become a mainstream approach for visual perception and is widely applied in a variety of downstream tasks. However, existing methods typically r…
Towards Generalizing Temporal Action Segmentation to Unseen Views
Emad Bahrami, Olga Zatsarynna, Gianpiero Francesca +1
While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen v…