1 paper
Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie +3
Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visually grounded inputs. This shift introduces…