9 papers
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
Rafal Kocielnik, Pengrui Han, Peiyang Song +5
Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work docum…
A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback
Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen +7
Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, assessing the quality of trainer f…
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
Firdavs Nasriddinov, Rafal Kocielnik, Anima Anandkumar +1
High-quality intraoperative feedback from a surgical trainer is pivotal for improving trainee performance and long-term skill acquisition. Automating natural, trainer-style feedbac…
Beyond Ethics: How Inclusive Innovation Drives Economic Returns in Medical AI
Balagopal Unnikrishnan, Ariel Guerra Adames, Amin Adibi +10
While ethical arguments for fairness in healthcare AI are well-established, the economic and strategic value of inclusive design remains underexplored. This perspective introduces…
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Pengrui Han, Rafal Kocielnik, Peiyang Song +4
Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems…
Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale
Rafal Kocielnik, Min Kim, Penphob +5
Detecting prosociality in text--communication intended to affirm, support, or improve others' behavior--is a novel and increasingly important challenge for trust and safety systems…