7 papers
Visual prompt engineering for video models
Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performan…
Physics-IQ Verified
Tim Rädsch, Yuki M Asano, Hilde Kuehne +4
Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a multitude of downstream tasks, including world modeling. To adva…
Unified Neural Scaling Laws
Ethan Caballero, Priyank Jaini, David Krueger +1
We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as mult…
Video models are zero-shot learners and reasoners
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6
The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models.…
Towards flexible perception with visual memory
Robert Geirhos, Priyank Jaini, Austin Stone +5
Training a neural network is a monolithic endeavor, akin to carving knowledge into stone: once the process is completed, editing the knowledge in a network is hard, since all infor…
Do generative video models understand physical principles?
Saman Motamed, Laura Culp, Kevin Swersky +2
AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world…