6 papers
STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos
Anshul Shah, Benjamin Lundell, Harpreet Sawhney +1
We address the problem of extracting key steps from unlabeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and per…
Self-supervised Learning with Local Contrastive Loss for Detection and Semantic Segmentation
Ashraful Islam, Ben Lundell, Harpreet Sawhney +3
We present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-lear…
Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition
Huseyin Coskun, Zeeshan Zia, Bugra Tekin +4
The lack of large-scale real datasets with annotations makes transfer learning a necessity for video activity understanding. We aim to develop an effective method for few-shot tran…
Video Analysis for Body-worn Cameras in Law Enforcement
Jason J. Corso, Alexandre Alahi, Kristen Grauman +4
The social conventions and expectations around the appropriate use of imaging and video has been transformed by the availability of video cameras in our pockets. The impact on law…
Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos
Mohamed Elhoseiny, Jingen Liu, Hui Cheng +2
We propose a new zero-shot Event Detection method by Multi-modal Distributional Semantic embedding of videos. Our model embeds object and action concepts as well as other available…
Depth Extraction from Videos Using Geometric Context and Occlusion Boundaries
S. Hussain Raza, Omar Javed, Aveek Das +3
We present an algorithm to estimate depth in dynamic video scenes. We propose to learn and infer depth in videos from appearance, motion, occlusion boundaries, and geometric contex…