Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 1 cited
NL-Eye: Abductive NLI for Images
Mor Ventura, Michael Toker, Nitay Calderon +3
Will a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor? Recent VLMs have demonstrated impressive capabilities, yet their ability to infer out…
cs.CV2024★ 1 cited
DOCCI: Descriptions of Connected and Contrasting Images
Yasumasa Onoe, Sunayana Rane, Zachary Berger +9
Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would al…
cs.CV2022
VASR: Visual Analogies of Situation Recognition
Yonatan Bitton, Ron Yosef, Eli Strugo +3
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Anal…