15 citations · 22 across the 11 of their papers we have counts for
1 paper · 1 filter
Pedro Freire, ChengCheng Tan, Adam Gleave +2
Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' represent…