activity
20242026
collaborators

7 papers

cs.AI2026

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Lewei Xu, Yihao Ding, Zihan Xu +5

Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…

cs.LG2026

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4

Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…

cs.AI2026

Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding

Yihao Ding, Qiang Sun, Puzhen Wu +3

Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…

cs.SE2025

DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral

Qiang Sun, Sirui Li, Tingting Bi +4

Acquiring structured data from domain-specific, image-based documents such as scanned reports is crucial for many downstream tasks but remains challenging due to document variabili…

cs.IR2025

SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration

Qiang Sun, Tingting Bi, Sirui Li +4

We present \textbf{SymbioticRAG}, a novel framework that fundamentally reimagines Retrieval-Augmented Generation~(RAG) systems by establishing a bidirectional learning relationship…

cs.LO2025

TimelineKGQA: A Comprehensive Question-Answer Pair Generator for Temporal Knowledge Graphs

Qiang Sun, Sirui Li, Du Huynh +2

Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and diff…