2 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demons…