9 papers
What Types of Human-AI Teams Exist?
Nathan Hughes, Ibrahim Habli
Human-AI teaming has received increasing attention in the literature. However, the range of studies conducted in multiple domains make it difficult to understand what types of team…
A Calibrated Memorization Index (MI) for Detecting Training Data Leakage in Generative MRI Models
Yash Deo, Yan Jia, Toni Lassila +5
Image generative models are known to duplicate images from the training data as part of their outputs, which can lead to privacy concerns when used for medical image generation. We…
WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
Zachary Ellis, Jared Joselowitz, Yash Deo +7
As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that st…
Evaluating Metrics for Safety with LLM-as-Judges
Kester Clegg, Richard Hawkins, Ibrahim Habli +1
LLMs (Large Language Models) are increasingly used in text processing pipelines to intelligently respond to a variety of inputs and generation tasks. This raises the possibility of…
Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems
Victoria J. Hodge, Colin Paterson, Ibrahim Habli
The operational capabilities and application domains of AI-enabled autonomous systems have expanded significantly in recent years due to advances in robotics and machine learning (…
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
Ernest Lim, Yajie Vera He, Jared Joselowitz +9
Despite the growing use of large language models (LLMs) in clinical dialogue systems, existing evaluations focus on task completion or fluency, offering little insight into the beh…