From the 1 of 9 linked papers with an AI index.
9 papers
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Lewei Xu, Yihao Ding, Zihan Xu +5
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…
MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
Yi Lin, Yihao Ding, Elana Benishay +8
MonteRET is an AI system that combines whole‑volume CT features with region‑level anatomical information and knowledge retrieval to automatically generate more complete and clinica…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…
MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation
Yi Lin, Yihao Ding, Yonghui Wu +1
Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language M…
STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System
Wenxiao Zhang, Yu Liu, Qiang sun +5
Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowledge graph construction require…
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
Biao Xiang, Soyeon Caren Han, Yihao Ding
Multi-hop question answering (QA) is widely used to evaluate the reasoning capabilities of large language models, yet most benchmarks focus on final answer correctness and overlook…