5 citations · 5 across the 1 of their papers we have counts for
1 paper · 1 filter
Runjin Chen, Andy Arditi, Henry Sleight +2
Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviat…