4 papers
Evaluating Metrics for Safety with LLM-as-Judges
Kester Clegg, Richard Hawkins, Ibrahim Habli +1
LLMs (Large Language Models) are increasingly used in text processing pipelines to intelligently respond to a variety of inputs and generation tasks. This raises the possibility of…
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
Ernest Lim, Yajie Vera He, Jared Joselowitz +9
Despite the growing use of large language models (LLMs) in clinical dialogue systems, existing evaluations focus on task completion or fluency, offering little insight into the beh…
Metrics that matter: Evaluating image quality metrics for medical image generation
Yash Deo, Yan Jia, Toni Lassila +5
Evaluating generative models for synthetic medical imaging is crucial yet challenging, especially given the high standards of fidelity, anatomical accuracy, and safety required for…
The case for delegated AI autonomy for Human AI teaming in healthcare
Yan Jia, Harriet Evans, Zoe Porter +5
In this paper we propose an advanced approach to integrating artificial intelligence (AI) into healthcare: autonomous decision support. This approach allows the AI algorithm to act…