3 papers
cs.CY2026
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
Alexandra DeLucia, Heyuan Huang, Sonal Joshi +3
LLM-as-a-Judge frameworks are increasingly trusted to automate evaluation in place of human experts, yet their reliability in high-stakes medical contexts remains unproven. We stre…
cs.CL2025
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
Alexandra DeLucia, Mark Dredze
Abstractive multi-document summarization (MDS) is the task of automatically summarizing information in multiple documents, from news articles to conversations with multiple speaker…
cs.DC2020
Analyzing HPC Support Tickets: Experience and Recommendations
Alexandra DeLucia, Elisabeth Moore
High performance computing (HPC) user support teams are the first line of defense against large-scale problems, as they are often the first to learn of problems reported by users.…