collaborators

11 papers

cs.CV2026

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Zhipeng Xu, Zulong Chen, Qing Liu +6

Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…

cs.IR2026

Memory Shot for Long-Term Dialogue

Chunyi Peng, Haidong Xin, Xuanshuo Sheng +7

Large Language Models (LLMs) have demonstrated strong capabilities in general conversation, instruction following, and complex reasoning. However, in long-term dialogue settings, t…

cs.CL2026

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

Denghao Ma, Qing Liu, Zulong Chen +5

Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings wit…

cs.CL2026

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

Xiang Feng, Jiawei Zhou, Zhangfeng Huang +6

Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer ac…

cs.CL2026

CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding

Zhipeng Xu, Junhao Ji, Yuqi Xiong +13

Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs acro…

cs.CV2026

UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents

Yifan Ji, Zhipeng Xu, Zhenghao Liu +7

Key Information Extraction (KIE) from real-world documents remains challenging due to substantial variations in layout structures, visual quality, and task-specific information req…