From the 1 of 10 linked papers with an AI index.
10 papers
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Runhan Shi, Quan Zhou, Yuqian Xu +14
The paper presents MedRealMM, a large-scale benchmark of real Chinese online medical consultations that includes both text and patient-uploaded images, and evaluates how well large…
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models
Jinru Ding, Chuchu Jiang, Lu Lu +12
Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinic…
Beyond Knowledge to Agency: Evaluating Expertise, Autonomy, and Integrity in Finance with CNFinBench
Jinru Ding, Chao Ding, Yidong Jiang +9
As large language models (LLMs) become high-privilege agents in risk-sensitive settings, they introduce systemic threats beyond hallucination, where minor compliance errors can cau…
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
Chao Ding, Mouxiao Bian, Tianbin Li +12
Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance requires auditable reasonin…
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation
Jiayong Zhu, Jiangtong Li, Jinru Ding +3
While Large Multimodal Models (LMMs) excel in general visual tasks, their deployment in specialized financial contexts remains insufficient. Existing benchmarks prioritize isolated…
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
Yiyun Zhu, Yidong Jiang, Ziwen Xu +4
Large language models (LLMs) are increasingly deployed in financial research workflows, where their role is evolving from single-model assistance for human analysts toward autonomo…