1 paper
Trent R Northen, Mingxun Wang
Large language models (LLMs) trained on internet-scale corpora can exhibit systematic biases that increase the probability of unwanted behavior. In this study, we examined potentia…