6 papers
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
Hamidreza Saffari, Francesco Pierri
Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across la…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
Mohammadamin Shafiei, Hamidreza Saffari
With the recent advances in Artificial Intelligence (AI) and Large Language Models (LLMs), the automation of daily tasks, like automatic writing, is getting more and more attention…
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
Mohammadamin Shafiei, Hamidreza Saffari, Nafise Sadat Moosavi
Large language models (LLMs) are known to be sensitive to input phrasing, but the mechanisms by which semantic cues shape reasoning remain poorly understood. We investigate this ph…
MultiHoax: A Dataset of Multi-hop False-Premise Questions
Mohammadamin Shafiei, Hamidreza Saffari, Nafise Sadat Moosavi
As Large Language Models are increasingly deployed in high-stakes domains, their ability to detect false assumptions and reason critically is crucial for ensuring reliable outputs.…
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
Hamidreza Saffari, Mohammadamin Shafiei, Donya Rooein +2
Creating globally inclusive AI systems demands datasets reflecting diverse social norms. Iran, with its unique cultural blend, offers an ideal case study, with Farsi adding linguis…