1 paper
An Lanji, Dawei Liu, Jin Li +3
Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhibit unfaithfulness: the stat…