7 papers
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
Parker Seegmiller, Sarah Masud Preum
LLMs are increasingly deployed in dynamic, real-world settings, where the distribution of user prompts can shift substantially over time as new tasks, prompts, and users are introd…
Medical Triage as Pairwise Ranking: A Benchmark for Urgency in Patient Portal Messages
Joseph Gatto, Parker Seegmiller, Timothy Burdick +4
Medical triage is the task of allocating medical resources and prioritizing patients based on medical need. This paper introduces the first large-scale public dataset for studying…
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
Parker Seegmiller, Joseph Gatto, Sarah E. Greer +4
Large language models (LLMs) show promise in drafting responses to patient portal messages, yet their integration into clinical workflows raises various concerns, including whether…
Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles
Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Khushnur Binte Jahangir +2
Large Language Models (LLMs) show impressive performance on many NLP benchmarks, yet their ability to reason in figurative, culturally grounded, and low-resource settings remains u…
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
Omar Sharif, Joseph Gatto, Madhusudan Basak +1
Event argument extraction identifies arguments for predefined event roles in text. Existing work evaluates this task with exact match (EM), where predicted arguments must align exa…
Socially Constructed Treatment Plans: Analyzing Online Peer Interactions to Understand How Patients Navigate Complex Medical Conditions
Madhusudan Basak, Omar Sharif, Jessica Hulsey +4
When faced with complex and uncertain medical conditions (e.g., cancer, mental health conditions, recovery from substance dependency), millions of patients seek online peer support…