3 papers
cs.CL2025
PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Farhan Farsi, Shayan Bali, Fatemeh Valeh +4
With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detec…
cs.CL2025
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
Mohammad Hosseini, Kimia Hosseini, Shayan Bali +2
Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Eval…
cs.CL2025
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
Farhan Farsi, Farnaz Aghababaloo, Shahriar Shariati Motlagh +8
As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While compre…