7 papers
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
Junzhe Huang, Xiaoxiao Sun, Yan Yang +6
Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…
STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System
Wenxiao Zhang, Yu Liu, Qiang sun +5
Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowledge graph construction require…
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
Yihao Ding, Qiang Sun, Puzhen Wu +3
Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…
DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral
Qiang Sun, Sirui Li, Tingting Bi +4
Acquiring structured data from domain-specific, image-based documents such as scanned reports is crucial for many downstream tasks but remains challenging due to document variabili…
SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration
Qiang Sun, Tingting Bi, Sirui Li +4
We present \textbf{SymbioticRAG}, a novel framework that fundamentally reimagines Retrieval-Augmented Generation~(RAG) systems by establishing a bidirectional learning relationship…