1 paper · 2 filters
Jingshen Zhang, Bo Wang, Yanlin Fu +4
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily…