1 paper
Han-Jun Ko, Jr-Jen Chen, Haobo Yuan +4
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallu…