50 citations · 69 across the 2 of their papers we have counts for
7 papers
TSP: Temporally-Sensitive Pretraining of Video Encoders for Localization Tasks
Humam Alwassel, Silvio Giancola, Bernard Ghanem
Due to the large memory footprint of untrimmed videos, current state-of-the-art video localization methods operate atop precomputed video clip features. These features are extracte…
Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar +3
Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other…
RefineLoc: Iterative Refinement for Weakly-Supervised Action Localization
Alejandro Pardo, Humam Alwassel, Fabian Caba Heilbron +2
Video action detectors are usually trained using datasets with fully-supervised temporal annotations. Building such datasets is an expensive task. To alleviate this problem, recent…
MortonNet: Self-Supervised Learning of Local Features in 3D Point Clouds
Ali Thabet, Humam Alwassel, Bernard Ghanem
We present a self-supervised task on point clouds, in order to learn meaningful point-wise features that encode local structure around each point. Our self-supervised network, name…
The ActivityNet Large-Scale Activity Recognition Challenge 2018 Summary
Bernard Ghanem, Juan Carlos Niebles, Cees Snoek +6
The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-…
Diagnosing Error in Temporal Action Detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia +1
Despite the recent progress in video understanding and the continuous rate of improvement in temporal action localization throughout the years, it is still unclear how far (or clos…