1 paper
Ummara Mumtaz, Aimen Noor, Awais Ahmed
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interv…