4 citations · 4 across the 1 of their papers we have counts for
1 paper · 1 filter
Farzan Karimi-Malekabadi, Suhaib Abdurahman, Zhivar Sourati +2
Socio-cognitive benchmarks for large language models (LLMs) often fail to predict real-world behavior, even when models achieve high benchmark scores. Prior work has attributed thi…