Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
Olivia Sturman, Aparna Joshi, Bhaktipriya Radharapu +2
Increasing use of large language models (LLMs) demand performant guardrails to ensure the safety of inputs and outputs of LLMs. When these safeguards are trained on imbalanced data…
cs.CL2024
ShieldGemma: Generative AI Content Moderation Based on Gemma
Wenjun Zeng, Yuchi Liu, Ryan Mullins +9
We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety r…