7 papers
Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations
Mohammad Amine Jradi, Faeze Ghorbanpour, Alexander Fraser
Hate speech annotation is costly, subjective, and prone to annotator disagreement, making large-scale dataset construction challenging. We systematically analyze how well large lan…
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
Ruhallah Niazi, Faeze Ghorbanpour, Alexander Fraser
Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We introduce PersLitEval, a bench…
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
Faeze Ghorbanpour, Alexander Fraser
Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs remains poorly understood. We p…
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
Faeze Ghorbanpour, Daryna Dementieva, Alexander Fraser
Considering the importance of detecting hateful language, labeled hate speech data is expensive and time-consuming to collect, particularly for low-resource languages. Prior work h…
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
Faeze Ghorbanpour, Daryna Dementieva, Alexander Fraser
Despite growing interest in automated hate speech detection, most existing approaches overlook the linguistic diversity of online content. Multilingual instruction-tuned large lang…
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
Lukas Edman, Helmut Schmid, Alexander Fraser
The CUTE benchmark showed that LLMs struggle with character understanding in English. We extend it to more languages with diverse scripts and writing systems, introducing EXECUTE.…