27 papers
Whareformer: Learning to Track What is Where in Long Egocentric Videos
Jacob Chalk, Saptarshi Sinha, Dima Damen +2
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining kno…
Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda +2
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts b…
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi +3
Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often st…
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Prajwal Gatti, Simon Jenni, Fabian Caba Heilbron +1
We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on d…
Improving and Evaluating Hand-Object Interaction Detection
Ahmad Darkhalil, Dima Damen, David Fouhey
Understanding hands and the objects they interact with, both directly and through tools, is a key step for tasks ranging from action perception to 3D reconstruction and robotics. O…
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
Zhifan Zhu, Yifei Huang, Yoichi Sato +1
Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we introduce the N-Body Problem: pre…