activity
20192022
most citedSpace-time Mixing Attention for Video Transformer

57 citations · 60 across the 6 of their papers we have counts for

collaborators

14 papers

cs.CV2022

REST: REtrieve & Self-Train for generative action recognition

Adrian Bulat, Enrique Sanchez, Brais Martinez +1

This work is on training a generative action/video recognition model whose output is a free-form action-specific caption describing the video (rather than an action class label). A…

cs.CV2022

SOS! Self-supervised Learning Over Sets Of Handled Objects In Egocentric Action Recognition

Victor Escorcia, Ricardo Guerrero, Xiatian Zhu +1

Learning an egocentric action recognition model from video data is challenging due to distractors (e.g., irrelevant objects) in the background. Further integrating object informati…

cs.CV2021

SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021

Swathikiran Sudhakaran, Adrian Bulat, Juan-Manuel Perez-Rua +5

This report presents the technical details of our submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021. To participate in the challenge we deployed spatio-temporal…

cs.CV202157 cited

Space-time Mixing Attention for Video Transformer

Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2

This paper is on video recognition using Transformers. Very recent attempts in this area have demonstrated promising results in terms of recognition accuracy, yet they have been al…

cs.CV2021

Few-shot Action Recognition with Prototype-centered Attentive Learning

Xiatian Zhu, Antoine Toisoul, Juan-Manuel Perez-Rua +3

Few-shot action recognition aims to recognize action classes with few training samples. Most existing methods adopt a meta-learning approach with episodic training. In each episode…

cs.CV20201 cited

Boundary-sensitive Pre-training for Temporal Localization in Videos

Mengmeng Xu, Juan-Manuel Perez-Rua, Victor Escorcia +5

Many video analysis tasks require temporal localization thus detection of content changes. However, most existing models developed for these tasks are pre-trained on general video…