1 paper · 1 filter
Trent R Northen, Mingxun Wang
Large language models (LLMs) trained on internet-scale corpora can exhibit systematic biases that increase the probability of unwanted behavior. In this study, we examined potentia…