1 citations · 3 across the 3 of their papers we have counts for
3 papers · 1 filter
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language s…
Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization
Edward Fish, Jon Weinbren, Andrew Gilbert
Temporal Action Localization (TAL) aims to identify actions' start, end, and class labels in untrimmed videos. While recent advancements using transformer networks and Feature Pyra…
MOFO: MOtion FOcused Self-Supervision for Video Understanding
Mona Ahmadian, Frank Guerin, Andrew Gilbert
Self-supervised learning (SSL) techniques have recently produced outstanding results in learning visual representations from unlabeled videos. Despite the importance of motion in s…