76 citations · 77 across the 3 of their papers we have counts for
10 papers
Accelerating scientific discovery with Co-Scientist
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin +48
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce C…
Reinforcement Learning Towards Broadly and Persistently Beneficial Models
Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab +5
As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especial…
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim +13
Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Pro…
A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
Peter Brodeur, Jacob M. Koshy, Anil Palepu +45
Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clin…
Complementary Human-AI Clinical Reasoning in Ophthalmology
Mertcan Sevgi, Fares Antaki, Abdullah Zafar Khan +26
Vision impairment and blindness are a major global health challenge where gaps in the ophthalmology workforce limit access to specialist care. We evaluate AMIE, a medically fine-tu…
Advancing Conversational Diagnostic AI with Multimodal Reasoning
Khaled Saab, Jan Freyberg, Chunjong Park +33
Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…