222 citations · 296 across the 3 of their papers we have counts for
1 paper · 1 filter
Oliver Klingefjord, Ryan Lowe, Joe Edelman
There is an emerging consensus that we need to align AI systems with human values (Gabriel, 2020; Ji et al., 2024), but it remains unclear how to apply this to language models in p…