2 papers
cs.CV2025
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
Mona Ahmadian, Amir Shirian, Frank Guerin +1
Real-world videos often contain overlapping events and complex temporal dependencies, making multimodal interaction modeling particularly challenging. We introduce DEL, a framework…
cs.CV2024
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language s…