12 citations · 14 across the 9 of their papers we have counts for
1 paper · 2 filters
Cassidy Langenfeld, Claas Beger, Gloria Geng +4
Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans…