Publications (10)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Framework
Joe Watson, Ivan O'Connor, Chia-Wen Chen +3
Psychological assessments commonly rely on rating-scale items, which require respondents to condense complex experiences into predefined categories. Although rich, unstructured tex…
Latent Human Traits in the Language of Social Media: An Open-Vocabulary Approach
Vivek Kulkarni, Margaret L. Kern, David Stillwell +5
Over the past century, personality theory and research has successfully identified core sets of characteristics that consistently describe and explain fundamental differences in th…
Evaluating General-Purpose AI with Psychometrics
Xiting Wang, Liming Jiang, Jose Hernandez-Orallo +4
Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their…
The Golden Rule as a Heuristic to Measure the Fairness of Texts Using Machine Learning
Ahmed Izzidien, David Stillwell
In this paper we present a natural language programming framework to consider how the fairness of acts can be measured. For the purposes of the paper, a fair act is defined as one…
Large Language Models show both individual and collective creativity comparable to humans
Luning Sun, Yuzhuo Yuan, Yuan Yao +6
Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to hum…