1.3k citations · 1.5k across the 28 of their papers we have counts for
4 papers · 2 filters
Turbo Training with Token Dropout
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is an efficient training method for video tasks. We make three contributions: (1) We propose Turbo training, a simple and versatile training paradigm fo…
Prompt Generation Networks for Input-Space Adaptation of Frozen Vision Transformers
Jochem Loedeman, Maarten C. Stol, Tengda Han +1
With the introduction of the transformer architecture in computer vision, increasing model scale has been demonstrated as a clear path to achieving performance and robustness gains…
Temporal Alignment Networks for Long-term Video
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is a temporal alignment network that ingests long term video sequences, and associated text sentences, in order to: (1) determine if a sentence is align…
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc +24
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Fl…