collaborators

10 papers

cs.CL2026

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri +2

When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreeme…

cs.CL2026

Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring

Hadi Mohammadi, Shihan Wang, Masoume M. Raeissi +1

The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority deci…

cs.CL2026

EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models

Hadi Mohammadi, Anastasia Giachanou, Robert A. Bagheri

We present EvalMORAAL, a transparent chain-of-thought (CoT) framework that uses two scoring methods (log-probabilities and direct ratings) plus a model-as-judge peer review to eval…

cs.CL2026

Exploring Cultural Variations in Moral Judgments with Large Language Models

Hadi Mohammadi, Ayoub Bagheri

Large Language Models (LLMs) have shown strong performance across many tasks, but their ability to capture culturally diverse moral values remains unclear. In this paper, we examin…

cs.CL2026

Explainability-Based Token Replacement on LLM-Generated Text

Hadi Mohammadi, Anastasia Giachanou, Daniel L. Oberski +1

Generative models, especially large language models (LLMs), have shown remarkable progress in producing text that appears human-like. However, they often exhibit patterns that make…

cs.CL2025

Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs

Hadi Mohammadi, Tamas Kozak, Anastasia Giachanou

Chain-of-thought (CoT) reasoning has emerged as a powerful technique for improving the problem-solving capabilities of large language models (LLMs), particularly for tasks requirin…