activity
20242026
collaborators

8 papers

cs.CV2026

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

Junzhe Huang, Xiaoxiao Sun, Yan Yang +6

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…

cs.LG2026

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4

Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…

cs.DB2026

STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System

Wenxiao Zhang, Yu Liu, Qiang sun +5

Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowledge graph construction require…

cs.AI2026

Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding

Yihao Ding, Qiang Sun, Puzhen Wu +3

Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…

cs.SE2025

DocSpiral: A Platform for Integrated Assistive Document Annotation through Human-in-the-Spiral

Qiang Sun, Sirui Li, Tingting Bi +4

Acquiring structured data from domain-specific, image-based documents such as scanned reports is crucial for many downstream tasks but remains challenging due to document variabili…

cs.IR2025

SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration

Qiang Sun, Tingting Bi, Sirui Li +4

We present \textbf{SymbioticRAG}, a novel framework that fundamentally reimagines Retrieval-Augmented Generation~(RAG) systems by establishing a bidirectional learning relationship…