1 paper
Hao Wang, Limeng Qiao, Chi Zhang +4
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos rem…