1 paper · 1 filter
Nan Wang, Zhiwei Jin, Chen Chen +1
Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine…