1 paper · 1 filter
Xingjian Tao, Yiwei Wang, Yujun Cai +3
While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, precise coordinate prediction remains a significant challenge, particularly as high-resolutio…