8 papers
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
Tathagata Raha, Clement Christophe, Nada Saadi +4
Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks. We adapt the Cross-Examination Framework (CEF) for a reference-free,…
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
Clément Christophe, Wadood Mohammed Abdul, Prateek Munjal +3
As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient sa…
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel +8
While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnect…
Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks
Prateek Munjal, Clement Christophe, Ronnie Rajan +1
Instruction finetuning is standard practice for improving LLM performance, yet it remains unclear whether it enhances reasoning or merely induces surface-level pattern matching. We…
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
Svetlana Maslenkova, Clement Christophe, Marco AF Pimentel +7
Large language models offer transformative potential for healthcare, yet their responsible and equitable development depends critically on a deeper understanding of how training da…
Gene42: Long-Range Genomic Foundation Model With Dense Attention
Kirill Vishniakov, Boulbaba Ben Amor, Engin Tekin +12
We introduce Gene42, a novel family of Genomic Foundation Models (GFMs) designed to manage context lengths of up to 192,000 base pairs (bp) at a single-nucleotide resolution. Gene4…