5 papers
Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues
Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov
As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficia…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
Mohammadamin Shafiei, Hamidreza Saffari, Nafise Sadat Moosavi
Large language models (LLMs) are known to be sensitive to input phrasing, but the mechanisms by which semantic cues shape reasoning remain poorly understood. We investigate this ph…
MultiHoax: A Dataset of Multi-hop False-Premise Questions
Mohammadamin Shafiei, Hamidreza Saffari, Nafise Sadat Moosavi
As Large Language Models are increasingly deployed in high-stakes domains, their ability to detect false assumptions and reason critically is crucial for ensuring reliable outputs.…
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
Hamidreza Saffari, Mohammadamin Shafiei, Donya Rooein +2
Creating globally inclusive AI systems demands datasets reflecting diverse social norms. Iran, with its unique cultural blend, offers an ideal case study, with Farsi adding linguis…