356 citations · 954 across the 15 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 5 cited
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demons…
cs.CV2023★ 3 cited
Video Language Planning
Yilun Du, Mengjiao Yang, Pete Florence +10
We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…