565 citations · 7.1k across the 210 of their papers we have counts for
1 paper · 2 filters
Xinyang Geng, Hao Liu, Lisa Lee +3
Building scalable models to learn from diverse, multimodal data remains an open challenge. For vision-language data, the dominant approaches are based on contrastive learning objec…