1 paper · 1 filter
Joël Roman Ky, Salah Ghamizi, Maxime Cordy
Vision-Language Models (VLMs) map complex visual inputs to semantic spaces, but interpreting the cross-modal reasoning of VLMs currently relies on post-hoc explainers evaluated via…