activity
20172023
most citedFast Multi-frame Stereo Scene Flow with Motion Segmentation

29 citations · 57 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2023

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…

cs.CV20224 cited

Background Mixup Data Augmentation for Hand and Object-in-Contact Detection

Koya Tango, Takehiko Ohkawa, Ryosuke Furuta +1

Detecting the positions of human hands and objects-in-contact (hand-object detection) in each video frame is vital for understanding human activities from videos. For training an o…

cs.CV20213 cited

Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction

Takuma Yagi, Md Tasnimul Hasan, Yoichi Sato

Every hand-object interaction begins with contact. Despite predicting the contact state between hands and objects is useful in understanding hand-object interactions, prior methods…

cs.CV20213 cited

EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2021: Team M3EM Technical Report

Lijin Yang, Yifei Huang, Yusuke Sugano +1

In this report, we describe the technical details of our submission to the 2021 EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition. Leveraging multip…

cs.CV2020

Towards Visually Explaining Video Understanding Networks with Perturbation

Zhenqiang Li, Weimin Wang, Zuoyue Li +2

''Making black box models explainable'' is a vital problem that accompanies the development of deep learning networks. For networks taking visual information as input, one basic bu…

cs.CV20194 cited

Manipulation-skill Assessment from Videos with Spatial Attention Network

Zhenqiang Li, Yifei Huang, Minjie Cai +1

Recent advances in computer vision have made it possible to automatically assess from videos the manipulation skills of humans in performing a task, which breeds many important app…