1 paper
Xinhai Hou, Shaoyuan Xu, Manan Biyani +4
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful…