activity
20242026
most citedDocument Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

4 citations · 4 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

25 papers · 1 filter

cs.CV2026

From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

Jutao Xiao, Yuan Qu, Dongsheng Ma +7

Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To qu…

cs.CV2026

MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing

Bangrui Xu, Ziyang Miao, Xuanhe Zhou +7

VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together wi…

cs.CV2026

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

Zhongying Deng, Cheng Tang, Ziyan Huang +124

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in…

cs.CV2026

Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning

Junyuan Gao, Jiahe Song, Jiang Wu +11

Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether…

cs.CV2025

DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM

Qintong Zhang, Junyuan Zhang, Zhifei Ren +8

Document parsing aims to transform unstructured PDF images into semi-structured data, facilitating the digitization and utilization of information in diverse domains. While vision…

cs.CV2025

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

Shengyuan Ding, Xinyu Fang, Ziyu Liu +10

Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to…