1 citations · 3 across the 11 of their papers we have counts for
1 paper · 1 filter
Benjamin Feuer, Ameya Joshi, Chinmay Hegde
Vision language (VL) models like CLIP are robust to natural distribution shifts, in part because CLIP learns on unstructured data using a technique called caption supervision; the…