1 paper · 1 filter
Zeyu Wang, Xinming Xu
Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this enco…