1 paper · 1 filter
Zhiheng Wang, Bo Peng, Lai Wei +1
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal o…