4 papers · 1 filter
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
Yuan Pu, Yewon Chang, Furong Jia +4
People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic d…
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
Raymond Xiong, Furong Jia, Lionel Wong +1
Patients are increasingly using large language models (LLMs) to seek answers to their healthcare-related questions. However, benchmarking efforts in LLMs for question answering oft…
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
Furong Jia, Yuan Pu, Finn Guo +1
Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We…
Diagnosing our datasets: How does my language model learn clinical information?
Furong Jia, David Sontag, Monica Agrawal
Large language models (LLMs) have performed well across various clinical natural language processing tasks, despite not being directly trained on electronic health record (EHR) dat…