1 paper
Zheng Jiang, Yiming Chen, Nan He +4
Recent multimodal large language models (MLLMs) support Thinking with Images, invoking visual tools such as zooming and cropping to inspect image regions during inference. Yet thes…