3 papers
cs.CY2026
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
Alexandra DeLucia, Heyuan Huang, Sonal Joshi +3
LLM-as-a-Judge frameworks are increasingly trusted to automate evaluation in place of human experts, yet their reliability in high-stakes medical contexts remains unproven. We stre…
cs.CL2025
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
Jonathan Liu, Haoling Qiu, Jonathan Lasko +3
Recent research has shown that hallucinations, omissions, and biases are prevalent in everyday use-cases of LLMs. However, chatbots used in medical contexts must provide consistent…
cs.IR2025
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
Orion Weller, Benjamin Chang, Eugene Yang +7
Retrieval systems generally focus on web-style queries that are short and underspecified. However, advances in language models have facilitated the nascent rise of retrieval models…