2 citations · 4 across the 5 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
Mokshit Surana, Archit Rathod, Akshaj Satishkumar
Large Language Models (LLMs) trained on web-scale corpora inherently absorb toxic patterns from their training data. This leads to toxic degeneration where even innocuous prompts c…
cs.LG2026
Fair and Calibrated Toxicity Detection with Robust Training and Abstention
Mokshit Surana
Fairness in toxicity classification involves three integrated axes: ranking, calibration, and abstention. Training-time interventions and post-hoc safety mechanisms cannot be evalu…