2 citations · 2 across the 1 of their papers we have counts for
1 paper
Renshan Zhang, Yibo Lyu, Rui Shao +3
Cropping high-resolution document images into multiple sub-images is the most widely used approach for current Multimodal Large Language Models (MLLMs) to do document understanding…