2 papers
cs.LG2025
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
Tom Sühr, Florian E. Dorner, Olawale Salaudeen +2
Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intel…
cs.CL2023
Challenging the Validity of Personality Tests for Large Language Models
Tom Sühr, Florian E. Dorner, Samira Samadi +1
With large language models (LLMs) like GPT-4 appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate personality traits…