5 papers
Theory Trace Card: Theory-Driven Socio-Cognitive Evaluation of LLMs
Farzan Karimi-Malekabadi, Suhaib Abdurahman, Zhivar Sourati +2
Socio-cognitive benchmarks for large language models (LLMs) often fail to predict real-world behavior, even when models achieve high benchmark scores. Prior work has attributed thi…
Realistic threat perception drives intergroup conflict: A causal, dynamic analysis using generative-agent simulations
Suhaib Abdurahman, Farzan Karimi-Malekabadi, Chenxiao Yu +2
Human conflict is often attributed to threats against material conditions and symbolic values, yet it remains unclear how they interact and which dominates. Progress is limited by…
Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions
Farzan Karimi-Malekabadi, Pooya Razavi, Sonya Powers
As educational systems evolve, ensuring that assessment items remain aligned with content standards is essential for maintaining fairness and instructional relevance. Traditional h…
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
Jackson Trager, Francielle Vargas, Diego Alves +6
Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks. Nevertheless, current evaluati…
Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking
Alireza S. Ziabari, Nona Ghazizadeh, Zhivar Sourati +3
Large Language Models (LLMs) exhibit impressive reasoning abilities, yet their reliance on structured step-by-step processing reveals a critical limitation. In contrast, human cogn…