Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Ajmal M., Abin Roy, Afthab Salam Kanniyan +4
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…
cs.CL2025
Red Teaming Large Language Models for Healthcare
Vahid Balazadeh, Michael Cooper, David Pellow +32
We present the design process and findings of the pre-conference workshop at the Machine Learning for Healthcare Conference (2024) entitled Red Teaming Large Language Models for He…