Publications (5)
PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Farhan Farsi, Shayan Bali, Fatemeh Valeh +4
With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detec…
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
Farhan Farsi, Shayan Bali, Mohammad Heydari Rad +2
The paper studies how multimodal large language models associate musical instruments with gender categories, creating a new dataset (Symphony-Bias) and finding that text modalities…
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
Farhan Farsi, Farnaz Aghababaloo, Shahriar Shariati Motlagh +8
As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While compre…
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
Mohammad Hosseini, Kimia Hosseini, Shayan Bali +2
Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Eval…
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
Mohammad Heydari Rad, Farhan Farsi, Shayan Bali +2
Nowadays, the usage of Large Language Models (LLMs) has increased, and LLMs have been used to generate texts in different languages and for different tasks. Additionally, due to th…