49 citations · 54 across the 2 of their papers we have counts for
6 papers · 1 filter
Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
Mandela Patrick, Dylan Campbell, Yuki M. Asano +5
In video transformers, the time dimension is often treated in the same way as the two spatial dimensions. However, in a scene where objects or the camera may move, a physical point…
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models
Po-Yao Huang, Mandela Patrick, Junjie Hu +3
This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focus on multilingual text-to-video search and propose a Transformer-based model tha…
Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning
Mandela Patrick, Yuki M. Asano, Bernie Huang +4
The quality of the image representations obtained from self-supervised learning depends strongly on the type of data augmentations used in the learning formulation. Recent papers h…
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano +4
The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to…
Labelling unlabelled videos from scratch with multi-modal self-supervision
Yuki M. Asano, Mandela Patrick, Christian Rupprecht +1
A large part of the current success of deep learning lies in the effectiveness of data -- more precisely: labelled data. Yet, labelling a dataset with human annotation continues to…
Understanding Deep Networks via Extremal Perturbations and Smooth Masks
Ruth Fong, Mandela Patrick, Andrea Vedaldi
The problem of attribution is concerned with identifying the parts of an input that are responsible for a model's output. An important family of attribution methods is based on mea…