40 papers
From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing
Jutao Xiao, Yuan Qu, Dongsheng Ma +7
Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To qu…
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
Haojie Hu, Chenhao Dang, Yaojia Liu +3
Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs;…
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation
Zhiyuan Zhao, Bin Wang, Linke Ouyang +5
In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation. Within each loop iteration, the MLLM-DataEngine…
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
Chenhao Dang, Dantong Zhu, Jun Yang +2
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle cross-modal…
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
Bangrui Xu, Ziyang Miao, Xuanhe Zhou +7
VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together wi…
MoDora: Tree-Based Semi-Structured Document Analysis System
Bangrui Xu, Qihang Yao, Zirui Tang +8
Semi-structured documents integrate diverse interleaved data elements (e.g., tables, charts, hierarchical paragraphs) arranged in various and often irregular layouts. These documen…