14 citations · 17 across the 4 of their papers we have counts for
1 paper · 1 filter
Hongyu Zhu, Sichu Liang, Wenwen Wang +5
With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual grounding--have shown great poten…