3 papers
cs.LG2026
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada +4
LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurem…
cs.HC2024
An Equitable Experience? How HCI Research Conceptualizes Accessibility of Virtual Reality in the Context of Disability
Kathrin Gerling, Anna-Lena Meiners, Louisa Schumm +5
Creating accessible Virtual Reality (VR) is an ongoing concern in the Human-Computer Interaction (HCI) research community. However, there is little reflection on how accessibility…
cs.CL2024
Measuring Psychological Depth in Language Models
Fabrice Harel-Canada, Hanyu Zhou, Sreya Muppalla +4
Evaluations of creative stories generated by large language models (LLMs) often focus on objective properties of the text, such as its style, coherence, and diversity. While these…