2.9k citations · 4.7k across the 52 of their papers we have counts for
127 papers
Weakly-supervised Fingerspelling Recognition in British Sign Language Videos
K R Prajwal, Hannah Bull, Liliane Momeni +3
The goal of this work is to detect and recognize sequences of letters signed using fingerspelling in British Sign Language (BSL). Previous fingerspelling recognition methods have n…
End-to-end Tracking with a Multi-query Transformer
Bruno Korbar, Andrew Zisserman
Multiple-object tracking (MOT) is a challenging task that requires simultaneous reasoning about location, appearance, and identity of the objects in the scene over time. Our aim in…
A Tri-Layer Plugin to Improve Occluded Detection
Guanqi Zhan, Weidi Xie, Andrew Zisserman
Detecting occluded objects still remains a challenge for state-of-the-art object detectors. The objective of this work is to improve the detection for such objects, and thereby imp…
Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors
Vladimir Iashin, Weidi Xie, Esa Rahtu +1
The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spati…
Turbo Training with Token Dropout
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is an efficient training method for video tasks. We make three contributions: (1) We propose Turbo training, a simple and versatile training paradigm fo…
Compressed Vision for Efficient Video Understanding
Olivia Wiles, Joao Carreira, Iain Barr +2
Experience and reasoning occur across multiple temporal scales: milliseconds, seconds, hours or days. The vast majority of computer vision research, however, still focuses on indiv…