activity
20162026
most citedEPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

24 citations · 69 across the 55 of their papers we have counts for

collaborators
Showing 2019Show all

10 papers · 1 filter

cs.CV2019

Action Modifiers: Learning from Adverbs in Instructional Videos

Hazel Doughty, Ivan Laptev, Walterio Mayol-Cuevas +1

We present a method to learn a representation for adverbs from instructional videos using weak supervision from the accompanying narrations. Key to our method is the fact that the…

cs.CV2019

Weakly-Supervised Completion Moment Detection using Temporal Attention

Farnoosh Heidarivincheh, Majid Mirmehdi, Dima Damen

Monitoring the progression of an action towards completion offers fine grained insight into the actor's behaviour. In this work, we target detecting the completion moment of action…

cs.CV20195 cited

Sit-to-Stand Analysis in the Wild using Silhouettes for Longitudinal Health Monitoring

Alessandro Masullo, Tilo Burghardt, Toby Perrett +2

We present the first fully automated Sit-to-Stand or Stand-to-Sit (StS) analysis framework for long-term monitoring of patients in free-living environments using video silhouettes.…

cs.CV2019

Retro-Actions: Learning 'Close' by Time-Reversing 'Open' Videos

Will Price, Dima Damen

We investigate video transforms that result in class-homogeneous label-transforms. These are video transforms that consistently maintain or modify the labels of all videos in each…

cs.CV2019

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman +1

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a ran…

cs.CV2019

Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings

Michael Wray, Diane Larlus, Gabriela Csurka +1

We address the problem of cross-modal fine-grained action retrieval between text and video. Cross-modal retrieval is commonly achieved through learning a shared embedding space, th…