1 citations · 1 across the 7 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
VistaHop: Benchmarking Long-Horizon Visual DeepSearch
Hang He, Chuhuai Yue, Chengqi Dong +6
Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evi…
cs.CV2025
Training Multi-Image Vision Agents via End2End Reinforcement Learning
Chengqi Dong, Chuhuai Yue, Hang He +7
Recent VLM-based agents aim to replicate OpenAI O3's "thinking with images" via tool use, yet most open-source methods restrict inputs to a single image, limiting their applicabili…