6 papers · 1 filter
Generative Data Augmentation for Skeleton Action Recognition
Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye +1
Skeleton-based human action recognition is a powerful approach for understanding human behaviour from pose data, but collecting large-scale, diverse, and well-annotated 3D skeleton…
Human-AI Divergence in Ego-centric Action Recognition under Spatial and Spatiotemporal Manipulations
Sadegh Rahmaniboldaji, Filip Rybansky, Quoc C. Vuong +3
Humans consistently outperform state-of-the-art AI models in action recognition, particularly in challenging real-world conditions involving low resolution, occlusion, and visual c…
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
Mona Ahmadian, Amir Shirian, Frank Guerin +1
Real-world videos often contain overlapping events and complex temporal dependencies, making multimodal interaction modeling particularly challenging. We introduce DEL, a framework…
Interpretable Action Recognition on Hard to Classify Actions
Anastasia Anichenko, Frank Guerin, Andrew Gilbert
We investigate a human-like interpretable model of video understanding. Humans recognise complex activities in video by recognising critical spatio-temporal relations among explici…
DEAR: Depth-Enhanced Action Recognition
Sadegh Rahmaniboldaji, Filip Rybansky, Quoc Vuong +2
Detecting actions in videos, particularly within cluttered scenes, poses significant challenges due to the limitations of 2D frame analysis from a camera perspective. Unlike human…
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language s…