2 papers
cs.CL2026
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
Naseem Machlovi, Maryam Saleki, Ruhul Amin +5
As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholdi…
cs.AI2025
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
Naseem Machlovi, Maryam Saleki, Innocent Ababio +1
As AI systems become more integrated into daily life, the need for safer and more reliable moderation has never been greater. Large Language Models (LLMs) have demonstrated remarka…