7 citations · 26 across the 34 of their papers we have counts for
Showing 2023 · cs.CVShow all
2 papers · 2 filters
cs.CV2023★ 3 cited
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak +2
Foundation models like CLIP are trained on hundreds of millions of samples and effortlessly generalize to new tasks and inputs. Out of the box, CLIP shows stellar zero-shot and few…
cs.CV2023
Visual Data-Type Understanding does not emerge from Scaling Vision-Language Models
Vishaal Udandarao, Max F. Burg, Samuel Albanie +1
Recent advances in the development of vision-language models (VLMs) are yielding remarkable success in recognizing visual semantic content, including impressive instances of compos…