1 paper
Baoliang Tian, Yuxuan Si, Jilong Wang +13
Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largel…