2 papers
cs.CL2025
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
Qinyue Zheng, Salman Abdullah, Sam Rawal +7
LLMs are bound to transform healthcare with advanced decision support and flexible chat assistants. However, LLMs are prone to generate inaccurate medical content. To ground LLMs i…
cs.HC2025
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris +3
Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering…