1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2025
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
Mona Ahmadian, Amir Shirian, Frank Guerin +1
Real-world videos often contain overlapping events and complex temporal dependencies, making multimodal interaction modeling particularly challenging. We introduce DEL, a framework…
cs.CV2024★ 1 cited
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language s…
cs.CV2023★ 1 cited
MOFO: MOtion FOcused Self-Supervision for Video Understanding
Mona Ahmadian, Frank Guerin, Andrew Gilbert
Self-supervised learning (SSL) techniques have recently produced outstanding results in learning visual representations from unlabeled videos. Despite the importance of motion in s…