15 citations · 19 across the 3 of their papers we have counts for
3 papers
EgoViT: Pyramid Video Transformer for Egocentric Action Recognition
Chenbin Pan, Zhiqi Zhang, Senem Velipasalar +1
Capturing interaction of hands with objects is important to autonomously detect human actions from egocentric videos. In this work, we present a pyramid video transformer with a dy…
Large-Scale Automatic Labeling of Video Events with Verbs Based on Event-Participant Interaction
Andrei Barbu, Alexander Bridge, Dan Coroian +12
We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that des…
Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill +15
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as…