1 paper
Runjin Chen, Andy Arditi, Henry Sleight +2
Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviat…