29 citations · 57 across the 7 of their papers we have counts for
10 papers · 1 filter
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…
Background Mixup Data Augmentation for Hand and Object-in-Contact Detection
Koya Tango, Takehiko Ohkawa, Ryosuke Furuta +1
Detecting the positions of human hands and objects-in-contact (hand-object detection) in each video frame is vital for understanding human activities from videos. For training an o…
Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Takuma Yagi, Md Tasnimul Hasan, Yoichi Sato
Every hand-object interaction begins with contact. Despite predicting the contact state between hands and objects is useful in understanding hand-object interactions, prior methods…
EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2021: Team M3EM Technical Report
Lijin Yang, Yifei Huang, Yusuke Sugano +1
In this report, we describe the technical details of our submission to the 2021 EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition. Leveraging multip…
Towards Visually Explaining Video Understanding Networks with Perturbation
Zhenqiang Li, Weimin Wang, Zuoyue Li +2
''Making black box models explainable'' is a vital problem that accompanies the development of deep learning networks. For networks taking visual information as input, one basic bu…
Manipulation-skill Assessment from Videos with Spatial Attention Network
Zhenqiang Li, Yifei Huang, Minjie Cai +1
Recent advances in computer vision have made it possible to automatically assess from videos the manipulation skills of humans in performing a task, which breeds many important app…