1 paper
Alexandra DeLucia, Heyuan Huang, Sonal Joshi +3
LLM-as-a-Judge frameworks are increasingly trusted to automate evaluation in place of human experts, yet their reliability in high-stakes medical contexts remains unproven. We stre…