3 papers
cs.CV2026
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
Yuchuan Wu, Minghan Zhuo, Teng Fu +3
Document understanding with multimodal large language models (MLLMs) requires not only accurate answers but also explicit, evidence-grounded reasoning, especially in high-stakes sc…
cs.CL2025
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
Haiyang Yu, Yuchuan Wu, Fan Shi +18
Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and unders…
cs.CV2025
UMIT: Unifying Medical Imaging Tasks via Vision-Language Models
Haiyang Yu, Siyang Yi, Ke Niu +2
With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to…