1k citations · 1k across the 1 of their papers we have counts for
1 paper
Christoph Schuhmann, Romain Beaumont, Richard Vencu +13
Groundbreaking language-vision architectures like CLIP and DALL-E proved the utility of training on large amounts of noisy image-text data, without relying on expensive accurate la…