1 paper
Phillip Y. Lee, Jihyeon Je, Chanho Park +3
We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environmen…