1 paper
Haoyu Huang, Jinfa Huang, Zhongwei Wan +3
Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. How…