3 papers
cs.CL2026
How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh +2
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper e…
cs.RO2026
When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure
Victor Lockwood, Hasan Mahmud, Mohammad Javad Khojasteh +2
Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We propose an \textit{inverted eva…
cs.CL2024
A Linguistic Comparison between Human and ChatGPT-Generated Conversations
Morgan Sandler, Hyesun Choung, Arun Ross +1
This study explores linguistic differences between human and LLM-generated dialogues, using 19.5K dialogues generated by ChatGPT-3.5 as a companion to the EmpathicDialogues dataset…