Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
Nikhil Khandekar, Qiao Jin, Guangzhi Xiong +14
As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answeri…
cs.CL2024
AgentMD: Empowering Language Agents for Risk Prediction with Large-Scale Clinical Tool Learning
Qiao Jin, Zhizheng Wang, Yifan Yang +8
Clinical calculators play a vital role in healthcare by offering accurate evidence-based predictions for various purposes such as prognosis. Nevertheless, their widespread utilizat…