5 papers
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
Amy Rouillard, Sitwala Mundia, Linda Camara +8
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an…
Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
Bruce A. Bassett, Amy Rouillard, Sitwala Mundia +8
Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data, particularly in low and mid…
What You Prompt is What You Get: Increasing Transparency of Prompting Using Prompt Cards
Amandine M. Caut, Beimnet Zenebe, Amy Rouillard +1
The rapid advancement and impressive capabilities of large language models (LLMs) have given rise to the field of prompt engineering, the practice of crafting inputs to guide LLMs…
Representing data in words: A context engineering approach
Amandine M. Caut, Amy Rouillard, Beimnet Zenebe +3
Large language models (LLMs) have demonstrated remarkable potential across a broad range of applications. However, producing reliable text that faithfully represents data remains a…
Automated Quantum Algorithm Design using a Domain-Specific Language
Amy Rouillard, Matt Lourens, Francesco Petruccione
We present a computational method to automatically design the n-qubit realisations of quantum algorithms. Our approach leverages a domain-specific language (DSL) that enables the c…