306 citations · 309 across the 5 of their papers we have counts for
1 paper · 1 filter
Wanfu Wang, Qipeng Huang, Guangquan Xue +2
Vision Language Models (VLMs) have recently achieved significant progress in bridging visual perception and linguistic reasoning. Recently, OpenAI o3 model introduced a zoom-in sea…