Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
Siming Yan, Min Bai, Weifeng Chen +3
By combining natural language understanding, generation capabilities, and breadth of knowledge of large language models with image perception, recent large vision language models (…
cs.CV2024
AffordanceLLM: Grounding Affordance from Vision Language Models
Shengyi Qian, Weifeng Chen, Min Bai +3
Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires th…