Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
Francesco Ortu, Zhijing Jin, Diego Doimo +1
Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…
cs.CV2026
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5
Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…