3 papers
cs.CL2026
BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
Jiageng Wu, Bowen Gu, Ren Zhou +14
Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on l…
cs.LG2025
SmartAlert: Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Lab Utilization Reduction
April S. Liang, Fatemeh Amrollahi, Yixing Jiang +21
Repetitive laboratory testing unlikely to yield clinically useful information is a common practice that burdens patients and increases healthcare costs. Education and feedback inte…
cs.LG2025
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Yixing Jiang, Kameron C. Black, Gloria Geng +4
Recent large language models (LLMs) have demonstrated significant advancements, particularly in their ability to serve as agents thereby surpassing their traditional role as chatbo…