25 citations · 61 across the 19 of their papers we have counts for
1 paper · 2 filters
Yang Zhou, Zixuan Huang, Sunzhu Li +10
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that r…