3.8k citations · 6.6k across the 73 of their papers we have counts for
4 papers · 1 filter
TRecViT: A Recurrent Video Transformer
Viorica Pătrăucean, Xu Owen He, Joseph Heyward +10
We propose a novel block for \emph{causal} video modelling. It relies on a time-space-channel factorisation with dedicated blocks for each dimension: gated linear recurrent units (…
Improving fine-grained understanding in image-text pre-training
Ioana Bica, Anastasija Ilić, Matthias Bauer +8
We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multi…
One-Step Diffusion Distillation via Deep Equilibrium Models
Zhengyang Geng, Ashwini Pokle, J. Zico Kolter
Diffusion models excel at producing high-quality samples but naively require hundreds of iterations, prompting multiple attempts to distill the generation process into a faster net…
Visual Interaction Networks
Nicholas Watters, Andrea Tacchetti, Theophane Weber +3
From just a glance, humans can make rich predictions about the future state of a wide range of physical systems. On the other hand, modern approaches from engineering, robotics, an…