15 citations · 21 across the 3 of their papers we have counts for
3 papers
Scalable Pre-training of Large Autoregressive Image Models
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai +5
This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language…
Data Filtering Networks
Alex Fang, Albin Madappally Jose, Amit Jain +3
Large training sets have become a cornerstone of machine learning and are the foundation for recent advances in language modeling and multimodal learning. While data curation for p…
On Robustness in Multimodal Learning
Brandon McKinzie, Joseph Cheng, Vaishaal Shankar +3
Multimodal learning is defined as learning over multiple heterogeneous input modalities such as video, audio, and text. In this work, we are concerned with understanding how models…