1 citations · 1 across the 6 of their papers we have counts for
1 paper · 2 filters
Francesco Ortu, Zhijing Jin, Diego Doimo +1
Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…