7 papers
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Lewei Xu, Yihao Ding, Zihan Xu +5
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
Yihao Ding, Qiang Sun, Puzhen Wu +3
Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…
DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral
Qiang Sun, Sirui Li, Tingting Bi +4
Acquiring structured data from domain-specific, image-based documents such as scanned reports is crucial for many downstream tasks but remains challenging due to document variabili…
SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration
Qiang Sun, Tingting Bi, Sirui Li +4
We present \textbf{SymbioticRAG}, a novel framework that fundamentally reimagines Retrieval-Augmented Generation~(RAG) systems by establishing a bidirectional learning relationship…
TimelineKGQA: A Comprehensive Question-Answer Pair Generator for Temporal Knowledge Graphs
Qiang Sun, Sirui Li, Du Huynh +2
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and diff…