2 citations · 3 across the 5 of their papers we have counts for
1 paper · 1 filter
Katarina Slama, Alexandra Souly, Dishank Bansal +3
Preference-driven behavior in LLMs may be a necessary precondition for AI misalignment such as sandbagging: models cannot strategically pursue misaligned goals unless their behavio…