From the 1 of 9 linked papers with an AI index.
4 papers · 1 filter
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Lewei Xu, Yihao Ding, Zihan Xu +5
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…
MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation
Yi Lin, Yihao Ding, Yonghui Wu +1
Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language M…
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
Dhita Putri Pratama, Soyeon Caren Han, Yihao Ding
Large Vision-Language Models (LVLMs) achieve strong performance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasonin…
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
Yihao Ding, Qiang Sun, Puzhen Wu +3
Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…