activity
20082022
most citedThe Kinetics Human Action Video Dataset

2.9k citations · 4.7k across the 52 of their papers we have counts for

collaborators

127 papers

cs.CV20223 cited

Weakly-supervised Fingerspelling Recognition in British Sign Language Videos

K R Prajwal, Hannah Bull, Liliane Momeni +3

The goal of this work is to detect and recognize sequences of letters signed using fingerspelling in British Sign Language (BSL). Previous fingerspelling recognition methods have n…

cs.CV20222 cited

End-to-end Tracking with a Multi-query Transformer

Bruno Korbar, Andrew Zisserman

Multiple-object tracking (MOT) is a challenging task that requires simultaneous reasoning about location, appearance, and identity of the objects in the scene over time. Our aim in…

cs.CV20223 cited

A Tri-Layer Plugin to Improve Occluded Detection

Guanqi Zhan, Weidi Xie, Andrew Zisserman

Detecting occluded objects still remains a challenge for state-of-the-art object detectors. The objective of this work is to improve the detection for such objects, and thereby imp…

cs.CV20222 cited

Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Vladimir Iashin, Weidi Xie, Esa Rahtu +1

The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spati…

cs.CV20224 cited

Turbo Training with Token Dropout

Tengda Han, Weidi Xie, Andrew Zisserman

The objective of this paper is an efficient training method for video tasks. We make three contributions: (1) We propose Turbo training, a simple and versatile training paradigm fo…

cs.CV2022

Compressed Vision for Efficient Video Understanding

Olivia Wiles, Joao Carreira, Iain Barr +2

Experience and reasoning occur across multiple temporal scales: milliseconds, seconds, hours or days. The vast majority of computer vision research, however, still focuses on indiv…