3 papers
cs.AI2026
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Harry Rogers, Sally Shiels, Ashley Tomlinson +5
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and co…
cs.LG2026
Identity-Free Deferral For Unseen Experts
Joshua Strong, Pramit Saha, Yasin Ibrahim +2
Learning to Defer (L2D) improves AI reliability in decision-critical environments by training AI to either make its own prediction or defer the decision to a human expert. A key ch…
cs.CL2025
Trustworthy and Practical AI for Healthcare: A Guided Deferral System with Large Language Models
Joshua Strong, Qianhui Men, Alison Noble
Large language models (LLMs) offer a valuable technology for various applications in healthcare. However, their tendency to hallucinate and the existing reliance on proprietary sys…