2 citations · 2 across the 1 of their papers we have counts for
1 paper
Miles Wang, Tom Dupré la Tour, Olivia Watkins +8
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…