4 citations · 14 across the 19 of their papers we have counts for
1 paper · 2 filters
Haoqing Wang, Xingrun Xing, Wei Xia +2
Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling…