Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
cs.AI2024
Large Language Models show both individual and collective creativity comparable to humans
Luning Sun, Yuzhuo Yuan, Yuan Yao +6
Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to hum…
cs.AI2023
Evaluating General-Purpose AI with Psychometrics
Xiting Wang, Liming Jiang, Jose Hernandez-Orallo +4
Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their…