Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents
Zhongheng Zhou, Yi Sun, Huiguo He +6
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and…
cs.AI2026
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
Hao Yan, Yuliang Liu, Xingchen Liu +5
Existing Multimodal Large Language Models (MLLMs) suffer from significant performance degradation on the long document understanding task as document length increases. This stems f…