1 paper · 1 filter
Cassidy Langenfeld, Claas Beger, Gloria Geng +4
Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans…