1 paper
Andrei Marian Feier, Veysel Kocaman, Yigit Gul +6
Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions c…