2 citations · 2 across the 7 of their papers we have counts for
1 paper · 1 filter
Xinmiao Huang, Qisong He, Zhenglin Huang +5
Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous n…