2 papers
cs.CL2026
Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
Muhammad Deedahwar Mazhar Qureshi, Sannaan Khan, Muhammad Atif Qureshi +1
Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and…
cs.LG2026
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Association Attacks
Sannaan Khan, Muhammad U. S. Khan
As machine learning models grow more influential and opaque, algorithmic fairness and explainability are critical for ensuring accountability. However, we demonstrate that these au…