activity
20182022
most citedSpace-time Mixing Attention for Video Transformer

57 citations · 62 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CV20222 cited

Sylph: A Hypernetwork Framework for Incremental Few-shot Object Detection

Li Yin, Juan M Perez-Rua, Kevin J Liang

We study the challenging incremental few-shot object detection (iFSD) setting. Recently, hypernetwork-based approaches have been studied in the context of continuous and finetune-f…

cs.CV2021

SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021

Swathikiran Sudhakaran, Adrian Bulat, Juan-Manuel Perez-Rua +5

This report presents the technical details of our submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021. To participate in the challenge we deployed spatio-temporal…

cs.CV202157 cited

Space-time Mixing Attention for Video Transformer

Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2

This paper is on video recognition using Transformers. Very recent attempts in this area have demonstrated promising results in terms of recognition accuracy, yet they have been al…

cs.CV20201 cited

Boundary-sensitive Pre-training for Temporal Localization in Videos

Mengmeng Xu, Juan-Manuel Perez-Rua, Victor Escorcia +5

Many video analysis tasks require temporal localization thus detection of content changes. However, most existing models developed for these tasks are pre-trained on general video…

cs.CV20202 cited

Egocentric Action Recognition by Video Attention and Temporal Context

Juan-Manuel Perez-Rua, Antoine Toisoul, Brais Martinez +4

We present the submission of Samsung AI Centre Cambridge to the CVPR2020 EPIC-Kitchens Action Recognition Challenge. In this challenge, action recognition is posed as the problem o…

cs.CV2020

Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention

Juan-Manuel Perez-Rua, Brais Martinez, Xiatian Zhu +3

Attentive video modeling is essential for action recognition in unconstrained videos due to their rich yet redundant information over space and time. However, introducing attention…