8 citations · 9 across the 13 of their papers we have counts for
Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
Alexandra DeLucia, Heyuan Huang, Sonal Joshi +3
LLM-as-a-Judge frameworks are increasingly trusted to automate evaluation in place of human experts, yet their reliability in high-stakes medical contexts remains unproven. We stre…
cs.CY2025
Using large language models to promote health equity
Emma Pierson, Divya Shanmugam, Rajiv Movva +12
Advances in large language models (LLMs) have driven an explosion of interest about their societal impacts. Much of the discourse around how they will impact social equity has been…