122 citations · 218 across the 7 of their papers we have counts for
1 paper · 1 filter
Pedro Freire, ChengCheng Tan, Adam Gleave +2
Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' represent…