1 paper
Yongxin Wang, Ruizhe Zhou, Yueling Tang +4
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…