1 paper · 1 filter
Yi-Cheng Wang, Chu-Song Chen
Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue by the limited context window.…