45 citations · 147 across the 70 of their papers we have counts for
12 papers · 2 filters
Perception Test 2023: A Summary of the First Challenge And Outcome
Joseph Heyward, João Carreira, Dima Damen +2
The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking st…
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
Zhifan Zhu, Dima Damen
We propose the task of Hand-Object Stable Grasp Reconstruction (HO-SGR), the reconstruction of frames during which the hand is stably holding the object. We first develop the stabl…
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
Tomáš Souček, Dima Damen, Michael Wray +2
We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describi…
Learning from One Continuous Video Stream
João Carreira, Michael King, Viorica Pătrăucean +9
We introduce a framework for online learning from a single continuous video stream -- the way people and animals learn, without mini-batches, data augmentation or shuffling. This p…
Centre Stage: Centricity-based Audio-Visual Temporal Action Detection
Hanyuan Wang, Majid Mirmehdi, Dima Damen +1
Previous one-stage action detection approaches have modelled temporal dependencies using only the visual modality. In this paper, we explore different strategies to incorporate the…
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…