5 papers
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
Fangxin Shang, Yuan Xia, Dalu Yang +2
Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structured information exchange in clinical system…
Hypothesis-Driven Skill Optimization for LLM Agents
Fangxin Shang, Yehui Yang
External skills can improve action-oriented LLM agents without changing model weights, but persistent skill updates are risky when they are distilled from sparse or noisy trajector…
Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models
Yibin Zhao, Fangxin Shang, Dingrui Yang +1
Table question answering requires models to recover semantic relations encoded implicitly by two-dimensional layout, merged cells, and hierarchical headers. Current pipelines typic…
FCMBench-Video: Benchmarking Document Video Intelligence
Runze Cui, Fangxin Shang, Yehui Yang +3
Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Comp…
FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications
Yehui Yang, Dalu Yang, Fangxin Shang +7
FCMBench is the first large-scale and privacy-compliant multimodal benchmark for real-world financial credit applications, covering tasks and robustness challenges from domain spec…