26 citations · 27 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Large Language Models as Misleading Assistants in Conversation
Betty Li Hou, Kejian Shi, Jason Phang +3
Large Language Models (LLMs) are able to provide assistance on a wide range of information-seeking tasks. However, model outputs may be misleading, whether unintentionally or in ca…
cs.CL2023★ 26 cited
Pretraining Language Models with Human Preferences
Tomasz Korbak, Kejian Shi, Angelica Chen +5
Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, persona…