2 papers
cs.AI2026
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
Bangbang Zhou, Hangdi Xing, Yifan Chen +8
Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks h…
cs.CV2025
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
Chuwei Luo, Guozhi Tang, Qi Zheng +5
Multi-modal document pre-trained models have proven to be very effective in a variety of visually-rich document understanding (VrDU) tasks. Though existing document pre-trained mod…