3 papers
cs.CV2026
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Zhipeng Xu, Zulong Chen, Qing Liu +6
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…
cs.CL2026
CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding
Zhipeng Xu, Junhao Ji, Yuqi Xiong +13
Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs acro…
cs.CL2025
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
Fanwei Zhu, Jinke Yu, Zulong Chen +6
Automated resume information extraction is critical for scaling talent acquisition, yet its real-world deployment faces three major challenges: the extreme heterogeneity of resume…