2 papers
cs.AI2026
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
Alireza S. Ziabari, Kat Ellis, Colleen Chan +1
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer…
cs.CL2025
Benchmarking Generative AI for Scoring Medical Student Interviews in Objective Structured Clinical Examinations (OSCEs)
Jadon Geathers, Yann Hicke, Colleen Chan +5
Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and p…