2 papers
cs.CL2026
CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding
Zhipeng Xu, Junhao Ji, Yuqi Xiong +13
Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs acro…
cs.CV2024
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
Yuxin Tian, Mouxing Yang, Yunfan Li +4
Recent studies applied Parameter Efficient Fine-Tuning techniques (PEFTs) to efficiently narrow the performance gap between pre-training and downstream. There are two important fac…