activity
20242026
most citedBeyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

3 citations · 3 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20263 cited

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Wenting Chen, Guo Yu, Yiu-Fai Cheung +5

Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However, concerns persist regarding the reliabi…

cs.CL2026

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

Jing Hao, Siyuan Dai, Yongxin Zhang +13

Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for s…

cs.CL2025

Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

Jie Liu, Wenxuan Wang, Zizhan Ma +7

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large…

cs.CL2025

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models

Meidan Ding, Jipeng Zhang, Wenxuan Wang +6

Multimodal large language models (MLLMs) hold significant potential in medical applications, including disease diagnosis and clinical decision-making. However, these tasks require…

cs.CL2024

A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Jie Liu, Wenxuan Wang, Yihang Su +8

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. Ho…