7 papers
Sparse Autoencoder Features for Classifications and Transferability
Jack Gallifant, Shan Chen, Kuleen Sasse +3
Sparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transpa…
A Field Guide to Deploying AI Agents in Clinical Practice
Jack Gallifant, Katherine C. Kellogg, Matt Butler +19
Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implemen…
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
Shan Chen, Pedro Moreira, Yuxin Xiao +6
Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bas…
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
João Matos, Shan Chen, Siena Placino +13
Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fai…
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
Shan Chen, Mingye Gao, Kuleen Sasse +6
Background: Large language models (LLMs) are trained to follow directions, but this introduces a vulnerability to blindly comply with user requests even if they generate wrong info…
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
Shan Chen, Jack Gallifant, Mingye Gao +12
Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in t…