26 citations · 26 across the 1 of their papers we have counts for
1 paper
Tomasz Korbak, Kejian Shi, Angelica Chen +5
Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, persona…