1 paper
Haoran Zhao, Soyeon Caren Han, Eduard Hovy
Multimodal language models are typically evaluated through external behavior: selecting the correct image--text match, rejecting unsupported captions, or answering visual queries c…