11 papers
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
Tathagata Raha, Clement Christophe, Nada Saadi +4
Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks. We adapt the Cross-Examination Framework (CEF) for a reference-free,…
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
Clément Christophe, Wadood Mohammed Abdul, Prateek Munjal +3
As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient sa…
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel +8
While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnect…
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
Svetlana Maslenkova, Clement Christophe, Marco AF Pimentel +7
Large language models offer transformative potential for healthcare, yet their responsible and equitable development depends critically on a deeper understanding of how training da…
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
Nada Saadi, Tathagata Raha, Clément Christophe +3
This paper investigates the challenges of developing large language models (LLMs) proficient in both multilingual understanding and medical knowledge. We demonstrate that simply tr…
Named Clinical Entity Recognition Benchmark
Wadood M Abdul, Marco AF Pimentel, Muhammad Umar Salman +6
This technical report introduces a Named Clinical Entity Recognition Benchmark for evaluating language models in healthcare, addressing the crucial natural language processing (NLP…