26 citations · 53 across the 6 of their papers we have counts for
3 papers · 1 filter
Taken out of context: On measuring situational awareness in LLMs
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni +5
We aim to better understand the emergence of `situational awareness' in large language models (LLMs). A model is situationally aware if it's aware that it's a model and can recogni…
Pretraining Language Models with Human Preferences
Tomasz Korbak, Kejian Shi, Angelica Chen +5
Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, persona…
Aligning Language Models with Preferences through f-divergence Minimization
Dongyoung Go, Tomasz Korbak, Germán Kruszewski +3
Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional…