activity
20172022
most citedSpace-time Mixing Attention for Video Transformer

57 citations · 97 across the 10 of their papers we have counts for

collaborators

14 papers

cs.CV20225 cited

Relevance-based Margin for Contrastively-trained Video Retrieval Models

Alex Falcon, Swathikiran Sudhakaran, Giuseppe Serra +2

Video retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries…

cs.CV2021

SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021

Swathikiran Sudhakaran, Adrian Bulat, Juan-Manuel Perez-Rua +5

This report presents the technical details of our submission to the EPIC-Kitchens-100 Action Recognition Challenge 2021. To participate in the challenge we deployed spatio-temporal…

cs.CV202157 cited

Space-time Mixing Attention for Video Transformer

Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2

This paper is on video recognition using Transformers. Very recent attempts in this area have demonstrated promising results in terms of recognition accuracy, yet they have been al…

cs.CV202120 cited

Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries

Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz

We present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-no…

cs.CV20201 cited

FBK-HUPBA Submission to the EPIC-Kitchens Action Recognition 2020 Challenge

Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz

In this report we describe the technical details of our submission to the EPIC-Kitchens Action Recognition 2020 Challenge. To participate in the challenge we deployed spatio-tempor…

cs.CV2019

Gate-Shift Networks for Video Action Recognition

Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz

Deep 3D CNNs for video action recognition are designed to learn powerful representations in the joint spatio-temporal feature space. In practice however, because of the large numbe…