1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Hang He, Chuhuai Yue, Chengqi Dong +6
Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evi…