1 paper
Alona Strugatski, Licol Zeinfeld, Jason Cooper +3
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs.…