1 citations · 2 across the 3 of their papers we have counts for
3 papers
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert
This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language s…
MOFO: MOtion FOcused Self-Supervision for Video Understanding
Mona Ahmadian, Frank Guerin, Andrew Gilbert
Self-supervised learning (SSL) techniques have recently produced outstanding results in learning visual representations from unlabeled videos. Despite the importance of motion in s…
Heterogeneous Graph Learning for Acoustic Event Classification
Amir Shirian, Mona Ahmadian, Krishna Somandepalli +1
Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous…