4 papers · 1 filter
Beyond Language Modeling: An Exploration of Multimodal Pretraining
Shengbang Tong, David Fan, John Nguyen +18
The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models r…
Interpreting Physics in Video World Models
Sonia Joseph, Quentin Garrido, Randall Balestriero +5
A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate p…
EvalGIM: A Library for Evaluating Generative Image Models
Melissa Hall, Oscar Mañas, Reyhane Askari-Hemmat +14
As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound…
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
Mazda Moayeri, Michael Rabbat, Mark Ibrahim +1
Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today's best…