1 paper
Jingshen Zhang, Bo Wang, Yanlin Fu +4
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily…