4 papers
Automatic Replication of LLM Mistakes in Medical Conversations
Oleksii Proniakin, Diego Fajardo, Ruslan Nazarenko +1
Large language models (LLMs) are increasingly evaluated in clinical settings using multi-dimensional rubrics which quantify reasoning quality, safety, and patient-centeredness. Yet…
Medical Interpretability and Knowledge Maps of Large Language Models
Razvan Marinescu, Victoria-Elisabeth Gruber, Diego Fajardo
We present a systematic study of medical-domain interpretability in Large Language Models (LLMs). We study how the LLMs both represent and process medical knowledge through four di…
A Women's Health Benchmark for Large Language Models
Victoria-Elisabeth Gruber, Razvan Marinescu, Diego Fajardo +13
As large language models (LLMs) become primary sources of health information for millions, their accuracy in women's health remains critically unexamined. We introduce the Women's…
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
Diego Fajardo V., Oleksii Proniakin, Victoria-Elisabeth Gruber +1
We present MedPI, a high-dimensional benchmark for evaluating large language models (LLMs) in patient-clinician conversations. Unlike single-turn question-answer (QA) benchmarks, M…