1 paper · 1 filter
Jaeyoo Park, Jin Young Choi, Jeonghyung Park +1
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effec…