1 paper
Jiaxi Huang, Dongxu Wu, Hanwei Zhu +4
The rapid advancement of Multi-modal Large Language Models (MLLMs) has expanded their capabilities beyond high-level vision tasks. Nevertheless, their potential for Document Image…