6 papers
Video Generation Models are General-Purpose Vision Learners
Letian Wang, Chuhan Zhang, Rishabh Kabra +9
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a genera…
TRecViT: A Recurrent Video Transformer
Viorica PÄtrÄucean, Xu Owen He, Joseph Heyward +10
We propose a novel block for \emph{causal} video modelling. It relies on a time-space-channel factorisation with dedicated blocks for each dimension: gated linear recurrent units (…
Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
Chuhan Zhang, Guillaume Le Moing, Skanda Koppula +11
Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simpl…
Scaling 4D Representations
João Carreira, Dilara Gokay, Michael King +32
Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x20…
Learning from Streaming Video with Orthogonal Gradients
Tengda Han, Dilara Gokay, Joseph Heyward +6
We address the challenge of representation learning from a continuous stream of video as input, in a self-supervised manner. This differs from the standard approaches to video lear…
From Image to Video: An Empirical Study of Diffusion Representations
Pedro Vélez, Luisa F. PolanÃa, Yi Yang +4
Diffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis. This success has sparked interest in leveraging their represe…