1 paper
Yin Wang, Haotian Hu, Jineng Han +4
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen u…