10 papers
Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri +2
When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreeme…
Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring
Hadi Mohammadi, Shihan Wang, Masoume M. Raeissi +1
The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority deci…
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
Hadi Mohammadi, Anastasia Giachanou, Robert A. Bagheri
We present EvalMORAAL, a transparent chain-of-thought (CoT) framework that uses two scoring methods (log-probabilities and direct ratings) plus a model-as-judge peer review to eval…
Exploring Cultural Variations in Moral Judgments with Large Language Models
Hadi Mohammadi, Ayoub Bagheri
Large Language Models (LLMs) have shown strong performance across many tasks, but their ability to capture culturally diverse moral values remains unclear. In this paper, we examin…
Explainability-Based Token Replacement on LLM-Generated Text
Hadi Mohammadi, Anastasia Giachanou, Daniel L. Oberski +1
Generative models, especially large language models (LLMs), have shown remarkable progress in producing text that appears human-like. However, they often exhibit patterns that make…
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
Hadi Mohammadi, Tamas Kozak, Anastasia Giachanou
Chain-of-thought (CoT) reasoning has emerged as a powerful technique for improving the problem-solving capabilities of large language models (LLMs), particularly for tasks requirin…