2 papers
cs.CL2026
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer +32
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant…
cs.CL2024
MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
Nikhil Khandekar, Qiao Jin, Guangzhi Xiong +14
As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answeri…